DBOS spike fitness report: per-family adoption recommendation¶
Spike: ADR-0009 Epic: PRD #72 Tracers evaluated: registry compatibility tracer (#74–#77), web-enrichment value tracer (#78–#84) This report: #85, ADR-0009 bullet 5's required written fitness report
Recommendation¶
| Family | Recommendation | Irreducible blocker reproduced? |
|---|---|---|
| Registry ingestion | Adopt | No |
| Web enrichment | Adopt | No |
Neither family reproduced any of the seven irreducible blockers ADR-0009 bullet 16 defines as the only valid rejection grounds (failure to recover safely in Job mode; domain corruption or unsafe watermark movement; unavoidable duplicate paid work or budget overspend; incompatible Supabase transaction/connection behavior; unacceptable CVR hot-path degradation; a requirement for runtime DDL or excessive privilege; inability to version, drain, and operate workflows safely). Every finding below is either positive evidence or a correctable gap, per bullet 16's framing — this report does not treat unmeasured performance/cost/connection-headroom criteria as rejection grounds, and instead lists them as ordered follow-up work (§5).
This report does not authorize production cutover, Regnskabsdata migration, broader scraper migration, Conductor, or retirement of Cloud Run Jobs/Scheduler, per ADR-0009 bullet 5 and this issue's acceptance criteria. A production-adoption ADR is a separate future decision.
How to read this report¶
A mid-epic pivot governs everything after #76: the user explicitly
directed skipping "elaborate testing" — live GCP deployment/promotion/demo
cycles, exhaustive per-alert live verification, and performance
benchmarking — in favor of implementing tickets with real automated test
coverage and treating live verification as separate, later work. This is
why the evidence below splits cleanly at #76: #74–#76 have live GCP
demonstration evidence against da-bu-data-collection-test; #77–#84 have
code-and-CI evidence only, explicitly by direction, not by omission. That
split is called out per criterion below rather than glossed over, and the
unmeasured criteria it produces (§5) are exactly the "inconclusive evidence
extends the affected tracer" case ADR-0009's own decision text anticipates,
not a report-blocking gap.
Every claim below links to the file, test, or PR that is its reproducible evidence, per this issue's first acceptance criterion.
1. Registry ingestion tracer (#74–#77)¶
Correctness and recovery¶
- Startup, stable identity,
ingestion_runsmapping, interruption recovery — live-demonstrated againstda-bu-data-collection-testand closed with evidence in #75. One workflow per(source, entity, run_type, logical_window)(src/da_business_data_collection/orchestration/registry.py:RegistryWorkflowIdentity); the same ID is stored oningestion_runs, so restarting a logical window reuses one domain run rather than creating a duplicate. - Domain-state integrity under replay —
tests/test_dbos_registry.py::test_dbos_and_existing_delta_have_the_same_domain_resultproves the DBOS and non-DBOS paths produce identical rows, counts, artifact hashes, and watermarks; a forced termination immediately before or after the coarse database commit converges on replay and never advances an unsafe watermark (same test file, termination tests). - Which expensive stages are skipped on recovery — one deliberately coarse step (scroll, immutable raw write, in-memory mapping, psycopg upsert) per ADR-0009 bullet 6's cheap-acquisition exception; interruption replays that whole convergent unit rather than resuming mid-step.
- Version lifecycle: drain and fork — live-demonstrated in #75
(delay → CD-blocked promotion → drain → fork, each step actually
exercised against the live project).
run_drain/DrainTimeoutError/DrainRecoveryError(registry.py),gate-dbos-registry-retirement(.github/workflows/cd.yml) blockingdbos_registry_image_tagpromotion while pending work remains.
Domain-state and security isolation¶
- Dedicated
dbos_registrysystem schema anddbos_registry_runtimePostgres role, reached through its ownDBOS_SYSTEM_DATABASE_URL; a privileged CI/CD migration step (not runtime DDL) creates/upgrades schema objects and grants the runtime role DML/execute only, matching ADR-0009 bullet 8 exactly. - CI proves this negatively, not just declares it:
tests/test_dbos_registry.pyassertsdbos_registry_runtimecan neither perform DDL nor read outside its own schema, against a real Postgres container.
Observability and retention¶
- Structured JSON events (
log_event, shared across both families —src/da_business_data_collection/orchestration/structured_logging.py) and Cloud Trace export (src/da_business_data_collection/orchestration/telemetry.py) — live validation in #76 found and fixed two real bugs no code review or CI test caught:BatchSpanProcessorsilently drops spans if the process exits before its export schedule fires (fixed by an explicitflush_cloud_tracein afinallyon every exit path), andtelemetry.googleapis.comrequires agcp.project_idOTel Resource attribute thex-goog-user-projectheader alone does not satisfy (telemetry._resource). Both fixes shipped and are the currently deployed behavior. The remaining per-alert live verification items from #76's plan (forced Job failure → Pub/Sub alert; over-age work → alert; drain-recovered-with-errors → distinguishable alert) were explicitly scoped down per the mid-epic pivot and were not live-verified — code and CI evidence only (overage_workflows,_raise_if_recovered_with_errorsinregistry.py, all unit-tested). - History retention (#77): a daily
dbos-registry-cleanupJob (infra/dbos_registry_cleanup_job.tf) deletesSUCCESShistories after 30 days andERROR/CANCELLED/MAX_RECOVERY_ATTEMPTS_EXCEEDEDafter 90, through the supportedDBOSClient.delete_workflowsAPI, querying onlycompleted_beforea cutoff so active work is never a candidate (registry.cleanup_old_histories, unit-tested). Not live-verified — #77 shipped code+CI only, by explicit user choice (AskUserQuestion: "Implement retention Job only, skip benchmarking").
Performance, connection behavior, and cost — unmeasured¶
#77's own scope explicitly excluded this. When asked how to handle #77 (which mixed real feature work with fitness/performance benchmarking), the user chose "Implement retention Job only, skip benchmarking" over the alternative of also measuring DBOS-vs-non-DBOS wall-time/memory and connection-pool headroom. No CVR wall-time/memory comparison, no measured Supabase connection-pool headroom under DBOS's session-mode requirement (ADR-0009 bullet 8), and no cost projection at current ingester volumes exist. This is the single largest evidence gap for this family and the top item in §5.
2. Web-enrichment tracer (#78–#84)¶
Every stage below shares one DBOS application
(da-business-web-enrichment, dbos_web_enrichment schema,
web_dbos_runtime role) and writes into one web_company_presence table
through web_find_company.build_datasource. No web stage has ever been
deployed live — only the campaign shell's service account and two
disabled maintenance Jobs exist in da-bu-data-collection-test
(infra/dbos_web_job.tf, enable_dbos_web_job stays false). Every claim
in this section is code-and-CI evidence; none of it has live-GCP
confirmation, consistent with the mid-epic pivot governing #77 onward.
Correctness and domain-state integrity¶
- Find-company (#78): CVR-grep judgment against an already-fetched page,
proven behaviorally equivalent to the sibling PoC's own offline e2e
fixture (CVR 41527080, real trimmed data from the PoC's own cache —
tests/fixtures/web_enrichment/resights_homepage.json,tests/test_find_company.py). - SERP (#80), browser (#81), LLM (#82) each judge conservatively: SERP
returns a lower-confidence candidate only from the top organic result;
browser stops at fetch+classify+persist, deliberately not judging
ownership (that stays #78's job, applied by a later caller); LLM's
verdict vocabulary (
own-site/third-party-about-companyaccept,unrelated/insufficientdon't) is ported as-is from the PoC.
Recovery and repeated-paid-work avoidance¶
This is the family's strongest evidence category — directly the "measured repeated-call avoidance" ADR-0009 bullet 5 requires, though measured by automated test, not live GCP replay:
- SERP (
web_serp.py):serp_operationsdedup ledger keyed by a stable hash of (company, prompt/extractor version, model, input); a retried or recovered submission reuses the existing provider task instead of paying twice (tests/test_web_serp.py::test_run_serp_search_reuses_an_existing_operation_instead_of_resubmitting). - LLM (
web_llm.py):llm_operationsdedup ledger stores the judged verdict directly (the call is synchronous, unlike SERP's submit-then-poll). A real bug was caught here by code review, not by the first-pass design: a cross-workflow dedup hit (same operation identity, differentpipeline/extractor_version→ differentworkflow_id) was still being charged, because DBOS's own exactly-once tracking on the budget-reservation call only protects same-workflow-ID replay, not two distinct workflow IDs sharing one operation identity. Fixed by checking for an existing operation before reserving, at the workflow level, with a regression test (tests/test_web_llm.py::test_run_llm_campaign_does_not_double_charge_two_workflows_sharing_one_operation). This is exactly the class of defect ADR-0009 bullet 16 treats as correctable, not a rejection ground — and it was caught and closed before merge. - Browser (
web_browser.py): no payment involved, so its step-level DBOS memoization alone is sufficient once a fetch succeeds; per-domain politeness (a genuinely new design — the PoC's ownmean_delayconfiguration is dead code, never read by any path the PoC calls) uses the same atomic-claim shape, later extracted into a function shared with LLM's provider/model rate control (web_find_company.try_claim_rate_limit_slot). - Per-company isolation — proven for every stage's campaign fan-out
(
web_campaign.run_campaign,run_serp_campaign,run_browser_campaign,run_llm_campaign): one company's contract failure, budget exhaustion, or blocked/terminal fetch is caught and recorded distinctly, never failing or replaying unrelated companies. #78's isolation fixture uses a genuinely invalid CVR (InvalidCvrError, real domain validation — CVR numbers are always 8-digit positive integers), not an artificial test hook.
Security isolation¶
Same shape and same evidentiary bar as the registry tracer: dedicated
dbos_web_enrichment schema, web_dbos_runtime role, CI provisions both
roles against the same Postgres container and asserts web_dbos_runtime
can neither DDL nor read dbos_registry.*
(tests/test_web_find_company.py::test_web_runtime_role_cannot_ddl_or_read_the_registry_schema).
Version lifecycle¶
83 ported the registry tracer's already-proven drain/fork shape onto¶
dbos_web_enrichment (web_lifecycle.py, web_retirement_gate.py,
gate-dbos-web-retirement). The gate ran live in CD on merge and
correctly reported "dbos-web-campaign Job not deployed yet; nothing to
retire" — a real, if minimal, live confirmation that the gate mechanism
itself executes correctly end to end. The drain/fork lifecycle itself
(actually draining a version with real pending work) has never been
exercised live, since no version has ever had live pending work to drain.
Observability and retention¶
84 generalized the registry tracer's Cloud Trace mechanism¶
(telemetry.configure_cloud_trace(..., service_name=...), one
implementation, two distinguishable applications) and ported the overage/
cleanup functions (web_lifecycle.overage_web_workflows,
cleanup_old_web_histories). New: a campaign whose companies include any
"failed" result now exits non-zero and logs
dbos_web_campaign_anomaly_detected, so scraper-error/budget anomalies
ride the same job_failure alert every other scheduled Job uses, without
touching per-company isolation. None of this has live GCP confirmation
— no enabled Job exists yet to run it against.
Performance, connection behavior, and cost — unmeasured¶
No wall-time/memory measurement exists for any web stage (less directly
comparable to a "hot path" than CVR's bulk ingestion, but still
unmeasured). No cost projection exists. One specific, concrete cost signal
worth flagging: crawl4ai (added in #81) pulled in a large transitive
dependency tree (numpy, scipy, openai, nltk, nearly 60 packages) purely to
support headless-browser fetching that has never been exercised live —
this is a real, currently-unweighed Docker image size and Cloud Run
cold-start cost, not yet measured against actual constraints.
3. Deployment status — explicit gap¶
Only the following exist live in da-bu-data-collection-test, all
disabled or dormant:
dbos-registry-spike— live, enabled, actively exercised (prodregistry ingestion is unaffected, per ADR-0009 bullet 4).dbos-registry-overage-check,dbos-registry-cleanup— live, enabled.dbos-web-campaign,dbos-web-overage-check,dbos-web-cleanup— service accounts and Job/Scheduler resources exist; all three Jobs are disabled (enable_dbos_web_job = false). No web-enrichment code has ever run against real Cloud Run, real Secret Manager credentials, or real GCP telemetry endpoints.
Every web-enrichment stage was deliberately built as a standalone,
directly-invokable capability, never wired into an actual
find-company → SERP/browser/LLM escalation flow — that integration
decision was explicitly deferred at every ticket (#78 through #84) as
outside that ticket's acceptance criteria, not forgotten. web_campaign_main.py
currently only runs the find-company stage against fixture data.
4. Residual risks and production-readiness gaps¶
Per this issue's acceptance criteria, stated explicitly rather than implied:
- No live GCP evidence for anything after #76. #77 through #84 are
code-and-CI only, by direction. The registry tracer's #76 live
validation already found two real bugs invisible to code review and
unit tests (span-flush timing, the
gcp.project_idattribute) — there is no equivalent live confirmation that the web tracer's telemetry, drain, overage, or cleanup mechanisms behave correctly under real GCP conditions, only that their logic is correct in isolation. - The web-enrichment stages are not wired together. Find-company,
SERP, browser, and LLM each work independently and are each fully
tested independently, but nothing in this repo currently chains them
into the PoC's actual
01-find-companyescalation flow (deterministic check → SERP → browser fetch → LLM corroboration). A production cutover would need that integration built and tested as its own unit of work. - No cost projection exists for either family. This is the ADR-0009 bullet 5 minimum-content item with the least evidence of any listed there.
- No connection-headroom measurement exists. DBOS's session-mode connection requirement (ADR-0009 bullet 8) is implemented correctly (verified by working code and CI-gated Postgres integration tests across both applications) but its behavior under realistic concurrent load against the shared Supabase pooler has never been measured.
crawl4ai's dependency weight is unweighed against actual Cloud Run image-size/cold-start constraints (§2, "Performance... — unmeasured").- One intermittent test flake remains uninvestigated:
tests/test_web_browser.py::test_run_browser_campaign_isolates_a_failing_company_from_unrelated_companiesfailed twice mid-epic when running the full suite under load (never in isolation, never in CI). Symptoms differed between occurrences, consistent with a timing/thread-pool-teardown race across many DBOSQueue-based tests sharing one pytest process, not a logic defect — but the root cause was never confirmed.
5. Follow-up work (ordered)¶
Per ADR-0009 bullet 16, these are correctable findings and unmeasured criteria, not rejection grounds:
- Measure CVR wall-time/memory, DBOS-path vs. non-DBOS-path, at representative volumes (registry; explicitly skipped in #77).
- Measure Supabase connection-pool headroom under DBOS's session-mode requirement, for both applications under realistic concurrent load.
- Produce a cost projection at current ingester volumes for the registry tracer, and a comparable projection for web-enrichment once §5.5 below gives it real traffic to project from.
- Complete #76's originally-planned live verification: forced Job
failure → Pub/Sub alert; over-age work → alert; a drain that recovers
an errored workflow → distinguishable alert; registry staleness →
correlated
ingestion_runsrow. (Explicitly scoped down mid-epic, not abandoned.) - Live-deploy the web-enrichment campaign (enable
dbos-web-campaign, run the two-stage credential rotation already documented indocs/reference/dbos-web-find-company-tracer.md) and exercise its drain/fork/overage/cleanup/telemetry mechanisms against real GCP, the same evidentiary bar #74–#76 already met for the registry tracer. - Decide and build the find-company → SERP/browser/LLM escalation wiring as its own unit of work, with its own acceptance criteria.
- Evaluate
crawl4ai's image-size/cold-start cost against Cloud Run's actual constraints once the web campaign is live-deployable; consider whether a lighter fetch strategy is warranted if the cost proves high. - Investigate the intermittent
test_web_browser.pycampaign-isolation flake to a confirmed root cause.
6. Scope of the future production-adoption ADR¶
A later ADR, informed by this report and by §5's completed follow-up work, would need to decide at minimum: whether to retire the existing Cloud Run Job + Scheduler pattern for CVR/Regnskabsdata in favor of DBOS Job-mode checkpointing, or run them side by side indefinitely; whether and how to migrate Regnskabsdata and any broader scraper workload onto the same primitives; whether DBOS Cloud/Conductor becomes warranted once self-hosted operational cost is measured (§5.2–5.3); and the concrete cutover sequencing, rollback plan, and monitoring bar for a live production migration. None of that is decided or authorized here.