Skip to content

Web enrichment deepens collection and evidence modules

Status: superseded by ADR-0029 Date: 2026-09-07 Decider: Chris Amends: ADR-0016 Preserves: ADR-0014 and the final amendments in ADR-0015

The maintainer accepted all six architecture-review recommendations and their related findings. The campaign ceiling stays at 30 companies by default and becomes configurable. Removing the ceiling was rejected. The requirements note supplements the existing domain and implementation; the interface design (removed with ADR-0029) compared three alternatives and recommends their concrete combination.

Accepted direction

  1. One page-acquisition module owns fast HTTP, browser escalation, usable content checks, pacing, budgets, deadlines, and attempt evidence. Full HTTP page artifacts follow the same retention rule as browser pages.
  2. One retained-evidence module owns unchanged raw source responses, append-only retrieval history, freshness selection, and projections. The final DataForSEO amendment in ADR-0015 already requires this behavior.
  3. Task-specific query planning precedes reusable collection of organic SERP, Google AI Mode, and website evidence. One verified-domain bundle remains distinct from third-party search evidence. Registry ingesters stay separate.
  4. Batch and ad hoc work use shared collection behavior with different execution priority. The configurable campaign cap defaults to 30 and refuses oversized requests; it does not silently split them. Priority must reach local execution as well as DataForSEO.
  5. A paid-operation module owns durable intent, short database transactions, uncertain provider acceptance, cost settlement, and recovery. An unknown acceptance must not cause a blind second purchase.
  6. The extraction module owns input coverage, exact evidence locators, validation, model identity, parsed-result storage, and independent outcomes. The configurable default LLM model is z-ai/glm-5.3-flash. Deterministic contact extraction and account-owned OpenRouter privacy policy remain.

Changes to earlier decisions

ADR-0016's browser-only acquisition wording is replaced by tiered acquisition. Its three resource classes remain; its fixed count of three physical queues must no longer prevent reserved ad hoc capacity. Work still commits with its outbox intent, and no process waits for DataForSEO completion. Normal dispatch and Cloud Tasks retries provide prompt recovery. The nightly operations job is the backstop for committed work that has no live task.

These changes do not permit a generic workflow engine, a runtime per small stage, multiple company domains in one bundle, automatic campaign replay, or a second people crawl. A new active-CVR request still returns the existing run identifier without creating work or cost. The proposed interface makes that conflict explicit instead of claiming the new intent was fulfilled.

Design status

Nightly residual recovery, 2026-09-11 (#760)

The one-minute repair trigger is removed. It called web-control about 1,440 times each day without application work and could keep the service resident. The accepted recovery interval is nightly.

Normal outbox dispatch still runs after each commit. Cloud Tasks still retries failed deliveries. The existing web-operations job runs the same bounded repair before broader reconciliation. It recovers undispatched outbox records, expired claims, and due unresolved paid operations. This amends the bounded one-minute interval accepted in this ADR and delivered by #646.

AI Mode competitor and customer relationships, 2026-09-10 (#714)

The competitor and customer goals use separate identity-specific AI Mode queries. Each query includes the full registry name and CVR number. The customer query asks for at most ten verified company names. The competitor query asks for at most twenty direct competitors by company name and CVR.

The retained AI Mode answer remains a source claim. Exact quote and locator validation is required before publication. A supported name enters the same company-mention ledger as website evidence. It becomes an observed edge only when the existing deterministic unique-name rule matches exactly one CVR company. Unknown and ambiguous names remain unresolved mentions. An empty validated answer is a successful explicit empty result. No second acquisition, extraction, or relationship pipeline is added.

Search-led office evidence, 2026-09-09 (#706)

The IGUS homepage and contact overview omit the office URL in both HTTP and rendered link metadata. Its content sitemap also omits the office URL. Public search returns the observed /headoffice page. More rendering or another contact navigation request did not provide the missing link in this probe.

The two-page navigation bound stays. When those pages do not establish identity, one persisted organic query restricted to the observed domain can supply one additional legal, contact, or office page. The query uses the full registry name and postcode. The existing collection module owns the purchase, receipt, retries, and admitted query budget. Acquisition owns the page budget, deadline, and retries. A delayed page is resumed under the same selection ID. There is no second ownership search, guessed URL, or domain-selection LLM call.

The live test exposed the storage boundary: acquisition cannot read search receipts. Control reads and verifies the retained receipt, excludes already selected pages, and persists one URL decision in the acquisition work payload. Acquisition receives the receipt identifier for provenance but does not read the search artifact. This preserves the existing storage permissions.

Search results supply a URL, not accepted identity evidence. A different domain is excluded. A conflicting labelled CVR on a checked page stops recovery. The approved exact-CVR or local registry-identity check remains mandatory, and non-CVR extraction remains limited to the matched company block. This amends the total ownership-page bound below with one measured search-led decision; it does not expand the navigation crawl or the global deadline.

Public HTTP and clear company ownership, 2026-09-09 (#699)

Chris approved public HTTP-only evidence and websites without a printed CVR when the website clearly belongs to the target company. This replaces the open policy questions below, not the strict TLS rule.

Explicit registry HTTP URLs keep their scheme. Email-domain leads and URLs without a scheme still start with HTTPS. If the first primary attempt fails with a certificate or connection error, acquisition can try the same host and path over HTTP once, within the existing attempt, pacing, and page deadline bounds. URLs with a query or explicit port do not receive this protocol change. Supporting pages use observed URLs; no guessed-path crawl is added. An HTTP redirect to broken HTTPS is not permission to disable verification. Each HTTP request has a fresh client without prior login state or cookies. This uses HTTPX's verified TLS default and client lifecycle, not a second insecure TLS client.

Accepted evidence, extraction reference menus, and exact source locators carry http-unencrypted or https-verified based on the actual response URL. The failed TLS attempt remains in the acquisition history. HTTP acceptance must never count as certificate repair or authenticated transport success.

The non-CVR rule requires the explicit registry website domain, full legal name, four-digit postcode, and normalized Danish registry phone. Name, postcode, and phone must occur in one local block of at most 600 characters, starting at the full legal name. A conflicting labelled CVR rejects the page. An email-domain guess, trading name, address alone, or LLM confidence cannot substitute for this evidence. The same identity gate applies to discovery planning; the LLM cannot override it. Registry verification retains its bound of two observed legal/contact pages, including an observed contact-to-office link. Contact and head-office links precede general terms. The known IGUS office page passes identity, but the live homepage route does not expose it within this bound. A third contact request did not recover it either, so the page limit is not increased on that evidence.

Without an exact CVR, the bundle freezes the registry identity used for the decision. Extractors see only the matched identity block, with offsets into the unchanged raw page. Other group products, contacts, social accounts, and people are not implicitly attributed to the target. Empty profile or people results are valid when that block states no such facts. This intentionally prefers limited, evidenced output to whole-group attribution.

registry-identity and ownership-unverified distinguish these decisions in run reports. The source-input recipe, ownership versions, extraction prompts, and page cache change with this policy. Old retained evidence is not deleted.

Strict browser transport, 2026-09-09 (#698, #700)

The certificate spike found that Crawl4AI 0.9.2 adds certificate-bypass launch flags independently of its context option. Browser acquisition now uses the existing Playwright dependency directly. The wrapper is removed, not retained as another path. HTTP and rendered pages share the same text/link projection. Each attempt has an isolated strict-TLS context; the acquisition module still owns pacing, redirects, retries, and the total deadline. Real browser tests cover invalid certificates, redirect targets, scripts, and proxy tunnels.

Website query planning uses the existing trading-name normalizer and postcode in one shared query. It does not add purchases or change ownership rules. Registry URLs keep their host, path, and query with HTTPS-first acquisition. HTTP-only evidence and non-CVR group-site attribution remain open in #699.

Coverage amendment, 2026-09-09 (#683–#688)

The measured run showed that first-link truncation discarded legal/contact links on a large product site. Registry ownership now checks at most two observed same-domain legal/contact pages when the homepage has no CVR match. An explicitly conflicting homepage CVR remains a rejection. A successful verification page is reused in the same bundle; there is no guessed-path crawl or relaxed ownership threshold. Legal/contact links have priority within the existing metadata bound. This amends ADR-0016's one-page registry ownership check, not its separate bound for people-page selection.

People and profile responses reached the 4096-token output limit on PFP's eight-page bundle. Those targets now admit 8192 tokens; source extraction stays at 4096. Actual target settings enter both extraction and paid-operation identity and provider capacity reservation. Closed schemas and compatible provider routing apply to all extraction calls. Incomplete results remain failures, never partial accepted JSON. Reasoning controls stay unchanged until completion metadata provides evidence for changing them.

Provider routing follows the OpenRouter structured-output contract. Its reasoning-token guidance explains why output-limit failures need reasoning-token measurements. The shared mail.tele.dk exclusion uses the provider-domain information in mySMTP's mail-provider status; it does not reject an explicit website supplied by the registry.

Reports separate resolved ownership from extractor success and nonempty output. Permanent certificate failures are not transient retries. Existing page and run deadlines stay unchanged while attempt and dispatch delays are measured.

The retained PFP retry then completed profile but reached the people cap again: 8192 completion tokens, of which 1801 were reasoning tokens, in 51.4 seconds. This supports a people-specific 16384-token output limit and a 120-second request limit, with at least 150 seconds for its paid claim. It does not support disabling reasoning or changing the successful profile/source settings. People responses remain bounded to 20 people and four work contacts per person, with distinct prompt and schema versions. The unchanged profile result remains reusable.

The next PFP response completed at 8980 output tokens but exposed an evidence isolation defect: one unsupported quote rejected the complete people array. People and work contacts now have independent exact-evidence checks. An invalid person is excluded; an invalid contact is excluded without discarding its grounded parent. Rejection counts remain visible. Schema and privacy failures still reject the result. This uses the same per-observation rule as source signals, without accepting partial JSON or weaker evidence (#695).

The linked design proposes method names, types, six physical queues, and detailed failure behavior. Those are reviewable interface and execution choices, not claims about deployed behavior. Code, configuration, data shapes, and release behavior are unchanged by this record. Implementation must replace obsolete paths rather than preserve parallel compatibility paths. Existing retained evidence must not be discarded as a side effect of removing runtime code.

Amendment 2026-09-11 — Patchright browser runtime (#747)

Browser acquisition now uses Patchright directly instead of Playwright. Patchright supplies the async browser API and the Chromium installation. This replaces the library named in the strict browser transport decision; it does not change the transport policy. Each attempt uses a fresh context, verifies certificates, and returns redirects to acquisition for validation. Service workers and automatic downloads remain disabled.

The change reduces known browser detection signals. The live acceptance comparison against the 14/29 blocked baseline remains a separate release check.