Web enrichment runs one company pass per task¶
Status: accepted Date: 2026-09-24 Decider: Chris Issue: #842 Supersedes: ADR-0016 and ADR-0020 Amends: ADR-0015 (website ownership without a printed CVR) Preserves: ADR-0014, the legal and privacy gates of ADR-0015, and the read contracts
Context¶
The goal of web enrichment is to collect useful information about a company with a CVR: its website, contacts, what it does, its people, its customers and competitors. The sources are search, Google AI Mode, and the company's own website.
The ADR-0016/0020 pipeline split one company into about twelve queued steps across three Cloud Run services and six queues, with an outbox, claims, leases, a minute repair loop, paid-operation receipts, and about 17,800 lines of web code. Its measured success rate was low: 34 of 274 run outcomes in the test environment resolved a website (12 %), and the last full campaign resolved 6 of 29 companies. A hand check of campaign 338 found these causes:
| Outcome | Companies | Cause |
|---|---|---|
fetch-blocked |
8 | 3 web firewalls refused the HTTP client (a normal browser request passes); 3 registry domains redirect to the company's new domain and a redirect was a failure; 2 dead hosts |
ownership-unverified |
8 | The registry names the website and it loads, but it does not print the CVR, and the name + postcode + phone block rule did not match (INCUBA A/S shows the registry phone and street) |
llm-rejected |
7 | Holding companies without a website; on the discovery path only a printed CVR could accept a site |
The machinery made each failure hard to see: the outcome names hid the cause, and one company's story was spread over work items, traces, and receipts.
Decision¶
One company pass¶
One process runs the whole pass for one company and waits for providers in place. DataForSEO live organic search (about 3 s, USD 0.002) and live AI Mode (about 6 s, USD 0.004) make an asynchronous task protocol unnecessary. The pass:
- reads the registry facts;
- collects leads: the registry website and email domain, the site AI Mode names as the website, then organic results (directories, social networks, and third-party profile pages are never leads);
- visits each lead's site: HTTP through the proxy with normal browser headers, one plain-HTTP try after an HTTPS transport failure, then a real browser through the proxy when the site blocks, needs JavaScript, or does not answer; redirects to any domain are followed;
- decides ownership from evidence (below);
- selects up to eight pages from the site's own links and sitemap, by topic quota: contact, about, people, offering, customers;
- stores every page's exact bytes and every search response before any reading;
- reads: deterministic contacts, social links, and technologies; one model call for profile, classifications, company mentions, and people; one model call for AI Mode competitors, customers, and profile statements;
- publishes everything in one transaction to the existing tables, then embeds the profile.
Each run keeps a readable log of every step and decision in
company_enrichment_run.log, with costs and timings.
Ownership from evidence¶
The registry facts a site shows decide the clear cases without a model:
| Evidence | Decision |
|---|---|
| Parked or for-sale page, or a registry directory page | refuse |
| The exact CVR on the site | accept |
| Another company's CVR and not ours | the model decides |
| A declared website (the registry website, or a site on its domain), and the site shows the name, phone, email, or street | accept |
| Another lead with the name and one of phone, email, street (not for holding companies) | accept |
| Another lead with no registry fact on the site, and a domain that is not the company's brand | refuse without a model call |
| Anything else | the model decides |
The model answers own, related, other, or unclear. own accepts. A
declared website is refused only for other (unrelated, parked, or a
directory): the company's own statement outweighs our doubt. Any other site
needs own, except a related group site whose domain is the company's brand
(gea.com for GEA PROCESS ENGINEERING A/S, nykredit.dk for NYKREDIT A/S). A related site's people are not
published: they may work for the group, not for this company.
This replaces the ADR-0015 rule that a site without a printed CVR needs the full legal name, postcode, and phone in one 600-character block. The remaining ADR-0015 gates stay: robots exclusion is ignored, TLS is always verified, raw pages keep full HTML, work contacts follow the attribution rules, and objections are applied before any write.
Every claim is quoted¶
The model quotes the page word for word for each claim. The code checks the quote against the page text (whitespace, letter case, and dash or quote-mark variants may differ) and stores its exact character range as the evidence locator. A claim without a matching quote is dropped; its siblings stay. People pass the existing people rules unchanged.
Execution¶
One Cloud Run service, web-enrichment, serves the request API and runs one
Cloud Tasks task per company run. Two queues separate batch work from ad hoc
requests. A run is claimed in Postgres, so a repeated delivery is a no-op.
An error releases the claim and returns HTTP 500 so the queue retries; the
third failed attempt ends the run as failed with the error named. A
scheduler calls a sweep every ten minutes to re-enqueue a run whose task was
lost or whose claim went stale.
A retried pass can repeat its searches and model calls. The exposure is about USD 0.01 per retry; the receipt and purchase-guard machinery that prevented it is removed. Paid results are not cached across runs: a refresh buys fresh evidence, and a re-extraction reads retained bytes and buys only model calls.
Outcomes an operator can act on¶
| Outcome | Reason | Meaning |
|---|---|---|
resolved |
cvr, registry-website, identity, judge-own, judge-related, judge-unclear, re-extract |
A verified website and how it was verified |
unresolved |
no-website |
Leads were checked; none belongs to the company. Search signals can still exist |
blocked |
site-unreachable |
A stated website (the registry website, or the site AI Mode calls the website) did not answer, or the registry email domain's server refused us; a 404 root or a mail-only domain is no-website |
failed |
search-failed, no-registry-row, no-retained-evidence, attempts-exhausted, error:<type> |
A system or provider fault |
Consequences¶
- Measured on the same 29 companies: 21 verified websites (old: 6), a mean of 28 s per company. On a 100-company shortlist of active operating companies that was not used to tune the rules: 73 have a website and an extracted profile, 71 company contacts, 71 company mentions, 46 named people, 93 AI Mode signals; a mean of 25 s per company (95th percentile 42 s); USD 0.006 search cost per company. The charter's M3 target (70 % with a website and a profile) is met. See the live acceptance record.
- The web code shrinks from about 17,800 to about 6,700 lines, the kept contact, people, mention, and technology rules included. The removed tables are the stage machine and paid-operation ledgers; the published tables and read functions are unchanged. Company mentions resolve only to active companies, through the trigram index instead of a scan of every company.
- Internal run and evidence tables get row-level security; the Supabase client roles could read and write them before.
- The request API keeps its shape.
goalsandfreshnessare removed; a request names CVRs and a scenario, and a re-extraction also names a reason. The default request cap is 500 companies. - Not built: directory adapters (ADR-0026), a second organic query, and a Facebook presence for companies without a website.