Skip to content

Web enrichment uses Postgres work state and Cloud Tasks across focused Cloud Run runtimes

Status: superseded by ADR-0029 Date: 2026-08-20 Deciders: Chris (solo founder) Supersedes: the web-enrichment half of ADR-0009 Depends on: ADR-0014 and ADR-0015 Tracking: issue #409

Context

The DBOS tracer proved the domain model, per-company failure isolation, content-addressed raw evidence, and paid-operation ledgers. Its production shape is weak: one Cloud Run Job owns browser processes, LLM calls, campaign coordination, and the submit-wait-collect protocol of DataForSEO. A measured run spent most of its wall time waiting after company work had parked, and one collector failure could stop unrelated work. The code also repeats stage-level campaign runners and performs extraction in different paths.

Decision

  1. Postgres is the source of truth for web-enrichment progress. A work mutation and its outbox record commit in one transaction. Cloud Tasks transports only a work ID and expected version. Handlers claim work idempotently and stale or duplicate deliveries are no-ops. DBOS is removed from web enrichment.
  2. A Cloud Run Job admits an on-demand campaign. Focused Cloud Run runtimes own control work, browser acquisition, LLM work, and the DataForSEO callback. Three Cloud Tasks queues isolate web-control, web-acquisition, and web-llm rate and resource limits. There is no queue for every small stage.
  3. Every Cloud Run HTTP runtime uses public ingress without a VPC. Cloud Tasks calls use IAM and OIDC authentication. The DataForSEO callback uses a single-use application token bound to the provider task ID. Each runtime has a dedicated least-privilege service account and only its required secrets.
  4. DataForSEO completion uses scheduled collection work. Submission stores the provider task ID and the charged cost, then creates a work item on web-control scheduled at the measured queue turnaround. Each attempt performs one tasks_ready request, collects whatever landed for this project, and either finishes the operation or reschedules itself with backoff until the stale-operation deadline. Cloud Tasks holds the delay, so no handler waits for another system, no resident worker exists, and no minutely dispatcher exists.

The accepted text had DataForSEO complete through an authenticated pingback into a callback runtime. That required allUsers on run.invoker for that service, because DataForSEO calls with no Google identity, and placing a load balancer, Cloud Armor or API Gateway in front does not remove the requirement from the backend. This project grants allUsers nowhere, so the pingback and the callback runtime are not built. The consequence is three new services rather than four, none with public ingress, and no single-use callback-token design. The DataForSEO stale-operation deadline becomes a value the measured run must set.

tasks_ready is a GET; a POST returns tasks: null, is indistinguishable from an empty queue, and has already cost two live campaigns their whole result set. 5. Campaign and task handlers try immediate outbox dispatch after commit. The existing nightly operations run repairs undispatched and stale work. There is no global every-minute dispatcher. 6. Candidate discovery first runs simple deterministic prechecks. If they do not produce a high-confidence candidate, DataForSEO supplies compact search evidence and an LLM selects one company domain. Every candidate still goes through primary-page ownership verification. The selector returns the domain root, primary URL, reasoning and confidence, or no_match. It never returns a ranked list of competing domains.

The accepted text also had the selector return useful same-domain URLs. No stage read them: page planning (decision 7) sees the same compact search evidence together with the fetched primary page, its links and its sitemap, so the selector's list carried nothing the planner did not already hold. The field and its prompt line are removed (#491) rather than kept for a requirement that does not exist. 7. Acquisition has two bounded phases. It first fetches the primary page. One LLM operation verifies ownership and selects useful same-domain links from search results, the page, and its sitemap. A second browser task fetches the selected pages sequentially and creates one evidence-bundle manifest. The system does not crawl or extract every raw page.

Amended 2026-08-24 (#154): people pages are selected deterministically, not by the planner, and carry their own bound. The planner's eight pages serve contacts, profile, products, cases, customers and mentions. Adding people to that prompt would make a team page compete with a product page for the same slot, so people evidence would arrive by displacing profile evidence. Instead a deterministic match runs over the same-domain link metadata and the bounded sitemap the pipeline already holds before planning: it scores URL slugs and anchor text against one Danish and English people-page vocabulary and takes at most two pages. The worst-case bundle is therefore ten sequential fetches, and the web-acquisition timeout must cover ten.

Amended 2026-08-25 (#509): the deterministic match also runs on the registry path. That path already holds the primary page's same-domain link metadata, but it does not fetch a sitemap or run the page planner. It matches only those primary-page links. When the match finds a people page, the existing bundle-acquisition stage fetches at most two selected pages before extraction. When it finds nothing, the one-page bundle still goes straight to extraction. Fetching a sitemap on every registry hit was rejected because it adds one request to the common path even when the primary page already exposes the useful links. The registry path therefore has a worst case of three pages; the discovery path keeps its worst case of ten.

The match itself costs nothing, because it reads only metadata that its path already holds. Probing likely paths (/kontakt, /om-os, /medarbejdere, …) was the alternative and is rejected: it spends one proxy fetch per guess and most guesses are 404. No probe fallback is built until a live run shows sites where the link match finds nothing. 8. One shared LLM operation path owns OpenRouter deduplication, budget, rate limits, retries, cost, and audit records. Separate versioned prompts and schemas perform domain selection, ownership and page planning, and bundle extraction. Contact, profile, mention, and people extractors have independent outcomes, and a retry runs only a missing or failed extractor version.

Amended 2026-08-24 (#154): the people extractor is the fourth member of the set, and it always runs. The sibling PoC made find-employees a second workflow with its own crawl and its own handoff from the first. This architecture has no workflows — it has one company enrichment run, one verified domain, one evidence bundle, and targeted extractors over that bundle. A second pipeline would fetch the same domain twice. So people extraction is an extractor, and the PoC's cross-workflow handoff requirement disappears with the second workflow that would have needed it.

Always-on rather than a third request kind: re-extract already runs one missing extractor version over a retained bundle, so a per-extractor request kind would be a second way to express what that one already expresses. The cost is one more LLM operation per refresh, per company. Amended 2026-09-11 (#743): requested bundle extractors are separate work items on the existing LLM queue. The bundle handoff creates one item per frozen target. Each target retains its extraction and provider identities, reservations, commits, and retries. Profile embedding belongs to its target. Completion is an idempotent readiness transition after required work becomes terminal. It creates one completion item when no required work remains; it does not poll goal work. Only the final transaction takes a per-run advisory lock, so independent provider calls overlap. Nightly repair retains responsibility for expired claims and failed dispatch. Cloud Tasks outbox calls use bounded groups of five, with individual success marks. First SERP collection uses the stored collect_after; a not-ready attempt starts the existing 30-second backoff.

  1. Refresh and re-extraction are different requests that converge on the same extraction module. Refresh creates current search, page, and bundle evidence. Re-extraction uses a retained bundle without DataForSEO or page fetches. Retention does not imply freshness.
  2. Postgres permits only one active web-enrichment run per CVR. A conflicting request returns the existing run ID and creates no tasks or provider cost. A future refresh after the active run is terminal is a new run, not shared work with an old campaign.
  3. A campaign completes when every company is terminal. Company outcomes are resolved, unresolved, blocked, budget-exhausted, or failed. A campaign with at least one failed company is completed-with-failures. Targeted retry does not replay the campaign.
  4. Cloud Tasks retries delivery failures, crashes, and timeouts. Expected transient provider failures are recorded in Postgres and create scheduled retry work through the outbox. Permanent failures become domain outcomes.
  5. The per-run wall clock measures elapsed time, including scheduled provider waits and page spacing. Tier m allows 3,600 seconds and tier l allows 7,200 seconds. The default tier therefore covers the bounded discovery path and several SERP collection backoffs without ending before extraction. These allowances are elapsed-time safety limits; page, query, and LLM-call budgets remain the cost limits.

Options considered

Option Decision
Keep all stages in one DBOS Cloud Run Job Rejected because unrelated resource profiles and external waiting share one process and failure boundary.
Create one Cloud Run runtime and queue per stage Rejected because it creates operational structure without a distinct resource or rate-limit boundary.
Use hardcoded rules as the main domain selector Rejected because measured matching quality is poor and the rules deepen with every exception.
Return several competing domains and try each Rejected because domain identity is one evidence decision. Extra pages must belong to the selected domain.
Poll DataForSEO every minute Rejected because no resident worker may exist only to wait. Scheduled collection work (decision 4) replaces the pingback: Cloud Tasks holds the delay between attempts, so the system still never runs a global every-minute poller.
Use Postgres work state, a transactional outbox, Cloud Tasks, and focused runtimes Chosen because it keeps durable state explicit and separates browser, LLM, and asynchronous provider behavior.

Consequences

  • The implementation must add explicit run, work, outbox, and evidence-bundle records while retaining the existing evidence, signal, budget, cost, and provider-operation ledgers that still express valid domain facts.
  • The final cutover deletes the DBOS web Job, workflows, queues, system schema, retirement gate, polling collector, and stage-only campaign runners. There is no dual runtime or backward-compatibility layer.
  • Runtime rate and concurrency settings can change independently without changing the domain flow.
  • The detailed target, flow, failure rules, and migration shape are in Web enrichment target architecture.