Web enrichment uses Postgres work state and Cloud Tasks across focused Cloud Run runtimes¶
Status: superseded by ADR-0029 Date: 2026-08-20 Deciders: Chris (solo founder) Supersedes: the web-enrichment half of ADR-0009 Depends on: ADR-0014 and ADR-0015 Tracking: issue #409
Context¶
The DBOS tracer proved the domain model, per-company failure isolation, content-addressed raw evidence, and paid-operation ledgers. Its production shape is weak: one Cloud Run Job owns browser processes, LLM calls, campaign coordination, and the submit-wait-collect protocol of DataForSEO. A measured run spent most of its wall time waiting after company work had parked, and one collector failure could stop unrelated work. The code also repeats stage-level campaign runners and performs extraction in different paths.
Decision¶
- Postgres is the source of truth for web-enrichment progress. A work mutation and its outbox record commit in one transaction. Cloud Tasks transports only a work ID and expected version. Handlers claim work idempotently and stale or duplicate deliveries are no-ops. DBOS is removed from web enrichment.
- A Cloud Run Job admits an on-demand campaign. Focused Cloud Run runtimes own
control work, browser acquisition, LLM work, and the DataForSEO callback.
Three Cloud Tasks queues isolate
web-control,web-acquisition, andweb-llmrate and resource limits. There is no queue for every small stage. - Every Cloud Run HTTP runtime uses public ingress without a VPC. Cloud Tasks calls use IAM and OIDC authentication. The DataForSEO callback uses a single-use application token bound to the provider task ID. Each runtime has a dedicated least-privilege service account and only its required secrets.
- DataForSEO completion uses scheduled collection work. Submission stores the
provider task ID and the charged cost, then creates a work item on
web-controlscheduled at the measured queue turnaround. Each attempt performs onetasks_readyrequest, collects whatever landed for this project, and either finishes the operation or reschedules itself with backoff until the stale-operation deadline. Cloud Tasks holds the delay, so no handler waits for another system, no resident worker exists, and no minutely dispatcher exists.
The accepted text had DataForSEO complete through an authenticated pingback
into a callback runtime. That required allUsers on run.invoker for that
service, because DataForSEO calls with no Google identity, and placing a
load balancer, Cloud Armor or API Gateway in front does not remove the
requirement from the backend. This project grants allUsers nowhere, so
the pingback and the callback runtime are not built. The consequence is
three new services rather than four, none with public ingress, and no
single-use callback-token design. The DataForSEO stale-operation deadline
becomes a value the measured run must set.
tasks_ready is a GET; a POST returns tasks: null, is
indistinguishable from an empty queue, and has already cost two live
campaigns their whole result set.
5. Campaign and task handlers try immediate outbox dispatch after commit. The
existing nightly operations run repairs undispatched and stale work. There
is no global every-minute dispatcher.
6. Candidate discovery first runs simple deterministic prechecks. If they do
not produce a high-confidence candidate, DataForSEO supplies compact search
evidence and an LLM selects one company domain. Every candidate still goes
through primary-page ownership verification. The selector returns the
domain root, primary URL, reasoning and confidence, or no_match. It never
returns a ranked list of competing domains.
The accepted text also had the selector return useful same-domain URLs. No stage read them: page planning (decision 7) sees the same compact search evidence together with the fetched primary page, its links and its sitemap, so the selector's list carried nothing the planner did not already hold. The field and its prompt line are removed (#491) rather than kept for a requirement that does not exist. 7. Acquisition has two bounded phases. It first fetches the primary page. One LLM operation verifies ownership and selects useful same-domain links from search results, the page, and its sitemap. A second browser task fetches the selected pages sequentially and creates one evidence-bundle manifest. The system does not crawl or extract every raw page.
Amended 2026-08-24 (#154): people pages are selected deterministically,
not by the planner, and carry their own bound. The planner's eight pages
serve contacts, profile, products, cases, customers and mentions. Adding
people to that prompt would make a team page compete with a product page
for the same slot, so people evidence would arrive by displacing profile
evidence. Instead a deterministic match runs over the same-domain link
metadata and the bounded sitemap the pipeline already holds before
planning: it scores URL slugs and anchor text against one Danish and
English people-page vocabulary and takes at most two pages. The
worst-case bundle is therefore ten sequential fetches, and the
web-acquisition timeout must cover ten.
Amended 2026-08-25 (#509): the deterministic match also runs on the registry path. That path already holds the primary page's same-domain link metadata, but it does not fetch a sitemap or run the page planner. It matches only those primary-page links. When the match finds a people page, the existing bundle-acquisition stage fetches at most two selected pages before extraction. When it finds nothing, the one-page bundle still goes straight to extraction. Fetching a sitemap on every registry hit was rejected because it adds one request to the common path even when the primary page already exposes the useful links. The registry path therefore has a worst case of three pages; the discovery path keeps its worst case of ten.
The match itself costs nothing, because it reads only metadata that its
path already holds. Probing likely paths (/kontakt, /om-os,
/medarbejdere, …) was the alternative and is rejected: it spends one
proxy fetch per guess and most guesses are 404. No probe fallback is built
until a live run shows sites where the link match finds nothing.
8. One shared LLM operation path owns OpenRouter deduplication, budget, rate
limits, retries, cost, and audit records. Separate versioned prompts and
schemas perform domain selection, ownership and page planning, and bundle
extraction. Contact, profile, mention, and people extractors have
independent outcomes, and a retry runs only a missing or failed extractor
version.
Amended 2026-08-24 (#154): the people extractor is the fourth member of the set, and it always runs. The sibling PoC made find-employees a second workflow with its own crawl and its own handoff from the first. This architecture has no workflows — it has one company enrichment run, one verified domain, one evidence bundle, and targeted extractors over that bundle. A second pipeline would fetch the same domain twice. So people extraction is an extractor, and the PoC's cross-workflow handoff requirement disappears with the second workflow that would have needed it.
Always-on rather than a third request kind: re-extract already runs one
missing extractor version over a retained bundle, so a per-extractor
request kind would be a second way to express what that one already
expresses. The cost is one more LLM operation per refresh, per company.
Amended 2026-09-11 (#743): requested bundle extractors are separate work
items on the existing LLM queue. The bundle handoff creates one item per
frozen target. Each target retains its extraction and provider identities,
reservations, commits, and retries. Profile embedding belongs to its target.
Completion is an idempotent readiness transition after required work becomes
terminal. It creates one completion item when no required work remains;
it does not poll goal work. Only the final transaction takes a per-run
advisory lock, so independent provider calls overlap. Nightly repair
retains responsibility for expired claims and failed dispatch. Cloud Tasks
outbox calls use bounded groups of five, with individual success marks.
First SERP collection uses the stored collect_after; a not-ready attempt
starts the existing 30-second backoff.
- Refresh and re-extraction are different requests that converge on the same extraction module. Refresh creates current search, page, and bundle evidence. Re-extraction uses a retained bundle without DataForSEO or page fetches. Retention does not imply freshness.
- Postgres permits only one active web-enrichment run per CVR. A conflicting request returns the existing run ID and creates no tasks or provider cost. A future refresh after the active run is terminal is a new run, not shared work with an old campaign.
- A campaign completes when every company is terminal. Company outcomes are
resolved,unresolved,blocked,budget-exhausted, orfailed. A campaign with at least one failed company iscompleted-with-failures. Targeted retry does not replay the campaign. - Cloud Tasks retries delivery failures, crashes, and timeouts. Expected transient provider failures are recorded in Postgres and create scheduled retry work through the outbox. Permanent failures become domain outcomes.
- The per-run wall clock measures elapsed time, including scheduled provider
waits and page spacing. Tier
mallows 3,600 seconds and tierlallows 7,200 seconds. The default tier therefore covers the bounded discovery path and several SERP collection backoffs without ending before extraction. These allowances are elapsed-time safety limits; page, query, and LLM-call budgets remain the cost limits.
Options considered¶
| Option | Decision |
|---|---|
| Keep all stages in one DBOS Cloud Run Job | Rejected because unrelated resource profiles and external waiting share one process and failure boundary. |
| Create one Cloud Run runtime and queue per stage | Rejected because it creates operational structure without a distinct resource or rate-limit boundary. |
| Use hardcoded rules as the main domain selector | Rejected because measured matching quality is poor and the rules deepen with every exception. |
| Return several competing domains and try each | Rejected because domain identity is one evidence decision. Extra pages must belong to the selected domain. |
| Poll DataForSEO every minute | Rejected because no resident worker may exist only to wait. Scheduled collection work (decision 4) replaces the pingback: Cloud Tasks holds the delay between attempts, so the system still never runs a global every-minute poller. |
| Use Postgres work state, a transactional outbox, Cloud Tasks, and focused runtimes | Chosen because it keeps durable state explicit and separates browser, LLM, and asynchronous provider behavior. |
Consequences¶
- The implementation must add explicit run, work, outbox, and evidence-bundle records while retaining the existing evidence, signal, budget, cost, and provider-operation ledgers that still express valid domain facts.
- The final cutover deletes the DBOS web Job, workflows, queues, system schema, retirement gate, polling collector, and stage-only campaign runners. There is no dual runtime or backward-compatibility layer.
- Runtime rate and concurrency settings can change independently without changing the domain flow.
- The detailed target, flow, failure rules, and migration shape are in Web enrichment target architecture.