Skip to content

DBOS spike fitness report: per-family adoption recommendation

Spike: ADR-0009 Epic: PRD #72 Tracers evaluated: registry compatibility tracer (#74–#77), web-enrichment value tracer (#78–#84) This report: #85, ADR-0009 bullet 5's required written fitness report

Recommendation

Family Recommendation Irreducible blocker reproduced?
Registry ingestion Adopt No
Web enrichment Adopt No

Neither family reproduced any of the seven irreducible blockers ADR-0009 bullet 16 defines as the only valid rejection grounds (failure to recover safely in Job mode; domain corruption or unsafe watermark movement; unavoidable duplicate paid work or budget overspend; incompatible Supabase transaction/connection behavior; unacceptable CVR hot-path degradation; a requirement for runtime DDL or excessive privilege; inability to version, drain, and operate workflows safely). Every finding below is either positive evidence or a correctable gap, per bullet 16's framing — this report does not treat unmeasured performance/cost/connection-headroom criteria as rejection grounds, and instead lists them as ordered follow-up work (§5).

This report does not authorize production cutover, Regnskabsdata migration, broader scraper migration, Conductor, or retirement of Cloud Run Jobs/Scheduler, per ADR-0009 bullet 5 and this issue's acceptance criteria. A production-adoption ADR is a separate future decision.

How to read this report

A mid-epic pivot governs everything after #76: the user explicitly directed skipping "elaborate testing" — live GCP deployment/promotion/demo cycles, exhaustive per-alert live verification, and performance benchmarking — in favor of implementing tickets with real automated test coverage and treating live verification as separate, later work. This is why the evidence below splits cleanly at #76: #74–#76 have live GCP demonstration evidence against da-bu-data-collection-test; #77–#84 have code-and-CI evidence only, explicitly by direction, not by omission. That split is called out per criterion below rather than glossed over, and the unmeasured criteria it produces (§5) are exactly the "inconclusive evidence extends the affected tracer" case ADR-0009's own decision text anticipates, not a report-blocking gap.

Every claim below links to the file, test, or PR that is its reproducible evidence, per this issue's first acceptance criterion.

1. Registry ingestion tracer (#74–#77)

Correctness and recovery

  • Startup, stable identity, ingestion_runs mapping, interruption recovery — live-demonstrated against da-bu-data-collection-test and closed with evidence in #75. One workflow per (source, entity, run_type, logical_window) (src/da_business_data_collection/orchestration/registry.py:RegistryWorkflowIdentity); the same ID is stored on ingestion_runs, so restarting a logical window reuses one domain run rather than creating a duplicate.
  • Domain-state integrity under replay — tests/test_dbos_registry.py::test_dbos_and_existing_delta_have_the_same_domain_result proves the DBOS and non-DBOS paths produce identical rows, counts, artifact hashes, and watermarks; a forced termination immediately before or after the coarse database commit converges on replay and never advances an unsafe watermark (same test file, termination tests).
  • Which expensive stages are skipped on recovery — one deliberately coarse step (scroll, immutable raw write, in-memory mapping, psycopg upsert) per ADR-0009 bullet 6's cheap-acquisition exception; interruption replays that whole convergent unit rather than resuming mid-step.
  • Version lifecycle: drain and fork — live-demonstrated in #75 (delay → CD-blocked promotion → drain → fork, each step actually exercised against the live project). run_drain/DrainTimeoutError/ DrainRecoveryError (registry.py), gate-dbos-registry-retirement (.github/workflows/cd.yml) blocking dbos_registry_image_tag promotion while pending work remains.

Domain-state and security isolation

  • Dedicated dbos_registry system schema and dbos_registry_runtime Postgres role, reached through its own DBOS_SYSTEM_DATABASE_URL; a privileged CI/CD migration step (not runtime DDL) creates/upgrades schema objects and grants the runtime role DML/execute only, matching ADR-0009 bullet 8 exactly.
  • CI proves this negatively, not just declares it: tests/test_dbos_registry.py asserts dbos_registry_runtime can neither perform DDL nor read outside its own schema, against a real Postgres container.

Observability and retention

  • Structured JSON events (log_event, shared across both families — src/da_business_data_collection/orchestration/structured_logging.py) and Cloud Trace export (src/da_business_data_collection/orchestration/telemetry.py) — live validation in #76 found and fixed two real bugs no code review or CI test caught: BatchSpanProcessor silently drops spans if the process exits before its export schedule fires (fixed by an explicit flush_cloud_trace in a finally on every exit path), and telemetry.googleapis.com requires a gcp.project_id OTel Resource attribute the x-goog-user-project header alone does not satisfy (telemetry._resource). Both fixes shipped and are the currently deployed behavior. The remaining per-alert live verification items from #76's plan (forced Job failure → Pub/Sub alert; over-age work → alert; drain-recovered-with-errors → distinguishable alert) were explicitly scoped down per the mid-epic pivot and were not live-verified — code and CI evidence only (overage_workflows, _raise_if_recovered_with_errors in registry.py, all unit-tested).
  • History retention (#77): a daily dbos-registry-cleanup Job (infra/dbos_registry_cleanup_job.tf) deletes SUCCESS histories after 30 days and ERROR/CANCELLED/MAX_RECOVERY_ATTEMPTS_EXCEEDED after 90, through the supported DBOSClient.delete_workflows API, querying only completed_before a cutoff so active work is never a candidate (registry.cleanup_old_histories, unit-tested). Not live-verified — #77 shipped code+CI only, by explicit user choice (AskUserQuestion: "Implement retention Job only, skip benchmarking").

Performance, connection behavior, and cost — unmeasured

#77's own scope explicitly excluded this. When asked how to handle #77 (which mixed real feature work with fitness/performance benchmarking), the user chose "Implement retention Job only, skip benchmarking" over the alternative of also measuring DBOS-vs-non-DBOS wall-time/memory and connection-pool headroom. No CVR wall-time/memory comparison, no measured Supabase connection-pool headroom under DBOS's session-mode requirement (ADR-0009 bullet 8), and no cost projection at current ingester volumes exist. This is the single largest evidence gap for this family and the top item in §5.

2. Web-enrichment tracer (#78–#84)

Every stage below shares one DBOS application (da-business-web-enrichment, dbos_web_enrichment schema, web_dbos_runtime role) and writes into one web_company_presence table through web_find_company.build_datasource. No web stage has ever been deployed live — only the campaign shell's service account and two disabled maintenance Jobs exist in da-bu-data-collection-test (infra/dbos_web_job.tf, enable_dbos_web_job stays false). Every claim in this section is code-and-CI evidence; none of it has live-GCP confirmation, consistent with the mid-epic pivot governing #77 onward.

Correctness and domain-state integrity

  • Find-company (#78): CVR-grep judgment against an already-fetched page, proven behaviorally equivalent to the sibling PoC's own offline e2e fixture (CVR 41527080, real trimmed data from the PoC's own cache — tests/fixtures/web_enrichment/resights_homepage.json, tests/test_find_company.py).
  • SERP (#80), browser (#81), LLM (#82) each judge conservatively: SERP returns a lower-confidence candidate only from the top organic result; browser stops at fetch+classify+persist, deliberately not judging ownership (that stays #78's job, applied by a later caller); LLM's verdict vocabulary (own-site/third-party-about-company accept, unrelated/insufficient don't) is ported as-is from the PoC.

Recovery and repeated-paid-work avoidance

This is the family's strongest evidence category — directly the "measured repeated-call avoidance" ADR-0009 bullet 5 requires, though measured by automated test, not live GCP replay:

  • SERP (web_serp.py): serp_operations dedup ledger keyed by a stable hash of (company, prompt/extractor version, model, input); a retried or recovered submission reuses the existing provider task instead of paying twice (tests/test_web_serp.py::test_run_serp_search_reuses_an_existing_operation_instead_of_resubmitting).
  • LLM (web_llm.py): llm_operations dedup ledger stores the judged verdict directly (the call is synchronous, unlike SERP's submit-then-poll). A real bug was caught here by code review, not by the first-pass design: a cross-workflow dedup hit (same operation identity, different pipeline/extractor_version → different workflow_id) was still being charged, because DBOS's own exactly-once tracking on the budget-reservation call only protects same-workflow-ID replay, not two distinct workflow IDs sharing one operation identity. Fixed by checking for an existing operation before reserving, at the workflow level, with a regression test (tests/test_web_llm.py::test_run_llm_campaign_does_not_double_charge_two_workflows_sharing_one_operation). This is exactly the class of defect ADR-0009 bullet 16 treats as correctable, not a rejection ground — and it was caught and closed before merge.
  • Browser (web_browser.py): no payment involved, so its step-level DBOS memoization alone is sufficient once a fetch succeeds; per-domain politeness (a genuinely new design — the PoC's own mean_delay configuration is dead code, never read by any path the PoC calls) uses the same atomic-claim shape, later extracted into a function shared with LLM's provider/model rate control (web_find_company.try_claim_rate_limit_slot).
  • Per-company isolation — proven for every stage's campaign fan-out (web_campaign.run_campaign, run_serp_campaign, run_browser_campaign, run_llm_campaign): one company's contract failure, budget exhaustion, or blocked/terminal fetch is caught and recorded distinctly, never failing or replaying unrelated companies. #78's isolation fixture uses a genuinely invalid CVR (InvalidCvrError, real domain validation — CVR numbers are always 8-digit positive integers), not an artificial test hook.

Security isolation

Same shape and same evidentiary bar as the registry tracer: dedicated dbos_web_enrichment schema, web_dbos_runtime role, CI provisions both roles against the same Postgres container and asserts web_dbos_runtime can neither DDL nor read dbos_registry.* (tests/test_web_find_company.py::test_web_runtime_role_cannot_ddl_or_read_the_registry_schema).

Version lifecycle

83 ported the registry tracer's already-proven drain/fork shape onto

dbos_web_enrichment (web_lifecycle.py, web_retirement_gate.py, gate-dbos-web-retirement). The gate ran live in CD on merge and correctly reported "dbos-web-campaign Job not deployed yet; nothing to retire" — a real, if minimal, live confirmation that the gate mechanism itself executes correctly end to end. The drain/fork lifecycle itself (actually draining a version with real pending work) has never been exercised live, since no version has ever had live pending work to drain.

Observability and retention

84 generalized the registry tracer's Cloud Trace mechanism

(telemetry.configure_cloud_trace(..., service_name=...), one implementation, two distinguishable applications) and ported the overage/ cleanup functions (web_lifecycle.overage_web_workflows, cleanup_old_web_histories). New: a campaign whose companies include any "failed" result now exits non-zero and logs dbos_web_campaign_anomaly_detected, so scraper-error/budget anomalies ride the same job_failure alert every other scheduled Job uses, without touching per-company isolation. None of this has live GCP confirmation — no enabled Job exists yet to run it against.

Performance, connection behavior, and cost — unmeasured

No wall-time/memory measurement exists for any web stage (less directly comparable to a "hot path" than CVR's bulk ingestion, but still unmeasured). No cost projection exists. One specific, concrete cost signal worth flagging: crawl4ai (added in #81) pulled in a large transitive dependency tree (numpy, scipy, openai, nltk, nearly 60 packages) purely to support headless-browser fetching that has never been exercised live — this is a real, currently-unweighed Docker image size and Cloud Run cold-start cost, not yet measured against actual constraints.

3. Deployment status — explicit gap

Only the following exist live in da-bu-data-collection-test, all disabled or dormant:

  • dbos-registry-spike — live, enabled, actively exercised (prod registry ingestion is unaffected, per ADR-0009 bullet 4).
  • dbos-registry-overage-check, dbos-registry-cleanup — live, enabled.
  • dbos-web-campaign, dbos-web-overage-check, dbos-web-cleanup — service accounts and Job/Scheduler resources exist; all three Jobs are disabled (enable_dbos_web_job = false). No web-enrichment code has ever run against real Cloud Run, real Secret Manager credentials, or real GCP telemetry endpoints.

Every web-enrichment stage was deliberately built as a standalone, directly-invokable capability, never wired into an actual find-company → SERP/browser/LLM escalation flow — that integration decision was explicitly deferred at every ticket (#78 through #84) as outside that ticket's acceptance criteria, not forgotten. web_campaign_main.py currently only runs the find-company stage against fixture data.

4. Residual risks and production-readiness gaps

Per this issue's acceptance criteria, stated explicitly rather than implied:

  1. No live GCP evidence for anything after #76. #77 through #84 are code-and-CI only, by direction. The registry tracer's #76 live validation already found two real bugs invisible to code review and unit tests (span-flush timing, the gcp.project_id attribute) — there is no equivalent live confirmation that the web tracer's telemetry, drain, overage, or cleanup mechanisms behave correctly under real GCP conditions, only that their logic is correct in isolation.
  2. The web-enrichment stages are not wired together. Find-company, SERP, browser, and LLM each work independently and are each fully tested independently, but nothing in this repo currently chains them into the PoC's actual 01-find-company escalation flow (deterministic check → SERP → browser fetch → LLM corroboration). A production cutover would need that integration built and tested as its own unit of work.
  3. No cost projection exists for either family. This is the ADR-0009 bullet 5 minimum-content item with the least evidence of any listed there.
  4. No connection-headroom measurement exists. DBOS's session-mode connection requirement (ADR-0009 bullet 8) is implemented correctly (verified by working code and CI-gated Postgres integration tests across both applications) but its behavior under realistic concurrent load against the shared Supabase pooler has never been measured.
  5. crawl4ai's dependency weight is unweighed against actual Cloud Run image-size/cold-start constraints (§2, "Performance... — unmeasured").
  6. One intermittent test flake remains uninvestigated: tests/test_web_browser.py::test_run_browser_campaign_isolates_a_failing_company_from_unrelated_companies failed twice mid-epic when running the full suite under load (never in isolation, never in CI). Symptoms differed between occurrences, consistent with a timing/thread-pool-teardown race across many DBOS Queue-based tests sharing one pytest process, not a logic defect — but the root cause was never confirmed.

5. Follow-up work (ordered)

Per ADR-0009 bullet 16, these are correctable findings and unmeasured criteria, not rejection grounds:

  1. Measure CVR wall-time/memory, DBOS-path vs. non-DBOS-path, at representative volumes (registry; explicitly skipped in #77).
  2. Measure Supabase connection-pool headroom under DBOS's session-mode requirement, for both applications under realistic concurrent load.
  3. Produce a cost projection at current ingester volumes for the registry tracer, and a comparable projection for web-enrichment once §5.5 below gives it real traffic to project from.
  4. Complete #76's originally-planned live verification: forced Job failure → Pub/Sub alert; over-age work → alert; a drain that recovers an errored workflow → distinguishable alert; registry staleness → correlated ingestion_runs row. (Explicitly scoped down mid-epic, not abandoned.)
  5. Live-deploy the web-enrichment campaign (enable dbos-web-campaign, run the two-stage credential rotation already documented in docs/reference/dbos-web-find-company-tracer.md) and exercise its drain/fork/overage/cleanup/telemetry mechanisms against real GCP, the same evidentiary bar #74–#76 already met for the registry tracer.
  6. Decide and build the find-company → SERP/browser/LLM escalation wiring as its own unit of work, with its own acceptance criteria.
  7. Evaluate crawl4ai's image-size/cold-start cost against Cloud Run's actual constraints once the web campaign is live-deployable; consider whether a lighter fetch strategy is warranted if the cost proves high.
  8. Investigate the intermittent test_web_browser.py campaign-isolation flake to a confirmed root cause.

6. Scope of the future production-adoption ADR

A later ADR, informed by this report and by §5's completed follow-up work, would need to decide at minimum: whether to retire the existing Cloud Run Job + Scheduler pattern for CVR/Regnskabsdata in favor of DBOS Job-mode checkpointing, or run them side by side indefinitely; whether and how to migrate Regnskabsdata and any broader scraper workload onto the same primitives; whether DBOS Cloud/Conductor becomes warranted once self-hosted operational cost is measured (§5.2–5.3); and the concrete cutover sequencing, rollback plan, and monitoring bar for a live production migration. None of that is decided or authorized here.