Skip to content

Web-enrichment adapters have explicit legal and privacy gates

Status: accepted Date: 2026-08-15 Deciders: Chris (solo founder) Depends on: ADR-0006 and ADR-0014 Current retention decision: browser page artifacts and minimized SERP artifacts in the raw store are kept for 365 days, not the 30 days below, as recorded in the 2026-08-19 amendment. Current page-artifact decision: browser-fetched page artifacts keep the full fetched Markdown and HTML without minimization, as recorded in the 2026-08-28 amendment. This supersedes the page-minimization rule below and in the first 2026-08-19 amendment. Current OpenRouter decision: endpoint data policy is configured on the OpenRouter account, not enforced per request in code, as recorded in the second 2026-08-19 amendment. Read the "OpenRouter" section below with that amendment. Current ownership decision: a website without a printed CVR belongs to the company when the registry facts on the site decide it, or the model judges it with those facts, as recorded in ADR-0029. This replaces the one-block name, postcode, and phone rule in the 2026-09-09 amendments of ADR-0020. Current work-contact decision: a work email is permitted on the same terms as a work phone, as recorded in the 2026-08-24 amendment inside "First-party professional person data". Every other rule in this ADR is unchanged.

Context

The web-enrichment pipeline can use DataForSEO search results, fetch public web pages with Crawl4AI, send evidence to an LLM through OpenRouter, and look up person phone numbers on Krak and De Gule Sider. Each adapter needs an explicit rule for source access, personal-data processing, and raw-artifact retention before production use.

The directory providers publish restrictions against automated collection, copying, caching, or storage. Public access to the pages does not remove this contract risk.

Decision

Krak and De Gule Sider

Direct automated access to public Krak and De Gule Sider pages is allowed. This decision knowingly accepts the risk that the access conflicts with the providers' published restrictions. It is not a finding that the access is legally permitted.

The adapter may retain a person's name, role, and work phone when the source publishes them in a company context. A directory result for the target company is sufficient context; it does not need an explicit job title or a separate person-to-phone statement. The adapter does not retain private-only listings or other personal directory data.

Robots exclusion

The browser adapter ignores robots.txt for every target. A robots denial does not produce a blocked outcome and does not stop a fetch. This is a deliberate source-access policy, not an operator override or an adapter error.

The browser may use normal browser actions and authentication or paid access that the controller lawfully owns. It must not use stolen credentials, exploit vulnerabilities, impersonate another person, or perform destructive access.

Automated CAPTCHA-solving services and browser-fingerprint evasion are allowed. This knowingly accepts the added contract and access-control risk. It does not relax the prohibitions on stolen credentials, vulnerability exploitation, identity deception, or destructive access.

A CAPTCHA provider does not need to promise zero retention or no training as a condition of use. The provider may therefore retain or reuse submitted content. This knowingly accepts an external data-sharing and privacy risk.

First-party professional person data

A company website may supply a person's name, job title, company role, work phone, work email, and evidence of a professional relationship. Processing rests on legitimate interest in building evidenced business relationships. A separate legitimate-interest assessment is optional and does not gate production use.

Private phone numbers, home addresses, CPR numbers, sensitive data, and personal facts that are not necessary for the professional relationship are not retained. Public availability alone does not make a personal fact necessary.

A phone number or an email address is a work contact point when it is published on a verified company-owned website or in a directory result for the target company. This company context is sufficient even when the source does not state a job title or an explicit person-to-value relationship. The rule can therefore classify an owner's personal mobile number as a work contact point when the company publishes or lists it in that context.

Amendment, 2026-08-24 (#154). The accepted text listed only the work phone, and prohibited "personal email addresses" without saying which addresses those are. That gap could not survive contact with a team page, which publishes work emails constantly. The decision is that the company context settles it, exactly as it already settles the phone: an address published on the verified company-owned domain is a work contact point, whatever the address itself looks like. A free-provider address a company publishes as its own contact is therefore retained. "Personal email address" now means an address obtained outside a verified company context, and that remains prohibited.

Company context alone does not attach a work contact point to a named person. When the source does not support that link through text, page structure, or structured data, the system stores a company work contact point and a separate person mention. It does not infer the relationship from proximity alone. An address whose own name matches the mention (jens.hansen@firma.dk) states the link in the value itself and is a person attribution; a generic address (info@firma.dk) states no person and stays a company contact point even when it sits beside one.

A person may object to storage of their work contact data. The data is then deleted, and a suppression record prevents collection from adding it again while the objection remains in effect.

Amendment, 2026-08-24 (#154). The accepted text called that record "minimized". It cannot be. To stop a value from returning, collection must recognise the value, so the record must hold something derived from the data it protects, and a plain hash of an eight-digit Danish phone number brute-forces in seconds. The suppression record therefore holds the objected value in clear, as decided in ADR-0008 decision 8. The consequence is stated rather than hidden: the suppression table is a readable list of the people who objected, it is personal data, and it carries the same row-level security as the rest of the hub.

A retained work contact point has no age-based expiry. It remains available indefinitely unless an objection or another explicit deletion rule applies. This knowingly accepts increased accuracy and storage-limitation risk; it is not a finding that indefinite retention is legally necessary or proportionate.

Later removal from the source does not delete or demote the value. The system continues to present it as current until the person objects or an operator removes it. This knowingly accepts the risk that a current read can contain an outdated or reassigned number or address.

Every displayed work contact point includes its source URL, collection time, and last verification time. Consumers can therefore see the age and origin even when an old value remains in the current read.

The controller does not proactively notify people whose work contact data is collected from websites or directories. This departs from the default GDPR Article 14 timing rule for indirect collection and knowingly accepts an unresolved compliance risk. It is not a finding that an Article 14 exception applies. A separate public privacy notice is optional and does not gate production use. Objections received through an available contact path are still handled under the deletion and suppression rule above.

Acceptance of this ADR is the legal and privacy gate for production use of the covered adapters. Existing implementation differences are follow-up work and do not keep production use disabled.

Fetched page text and HTML

Fetched page text and HTML are processed in memory and minimized before they enter the raw store. The retained artifact may contain the professional person data allowed above, but it must not contain the prohibited personal data. This source-specific minimized artifact follows ADR-0006 even though it is not a byte-for-byte copy of the upstream page.

Minimized first-party and directory page artifacts are retained for 30 days from retrieval and are then deleted. Derived evidence and domain records follow their separate retention rules.

Screenshots

A named extractor may capture and process a screenshot when the required information is best extracted from rendered pixels instead of page text or HTML. The screenshot is transient and is never stored. Only minimized derived evidence may enter the raw store or domain tables.

OpenRouter

Every OpenRouter request must select only provider endpoints that have zero data retention and do not train on prompts or responses. The operation fails closed when no compliant endpoint is available. It must not fall back to an endpoint with weaker data handling.

Full OpenRouter requests and raw provider responses are not stored. The audit record keeps only the input-artifact hash, prompt and model versions, minimized parsed result, token use, cost, and evidence provenance. An unreadable provider response is not persisted as a diagnostic artifact.

The minimized parsed-result cache is retained for 365 days. Its identity includes the model, prompt version, and input hash, so changed inputs or prompts do not reuse an older result.

OpenRouter prompt logging is disabled. The zero-data-retention and no-training rules apply to the selected model endpoint as well as to OpenRouter itself.

DataForSEO

The DataForSEO adapter may use the paid API to search for companies and their websites. It does not query for private persons. SERP results are minimized before storage under the same personal-data rules as fetched pages.

Minimized SERP payloads are retained in the raw store for 30 days from retrieval and are then deleted.

Retention summary

Artifact or record Retention
Minimized first-party page text and HTML 30 days
Minimized directory page text and HTML 30 days
Minimized DataForSEO SERP payload 30 days
Screenshot Never stored
Full OpenRouter request or raw response Never stored
Minimized OpenRouter parsed-result cache 365 days
Work phone Indefinite, until objection or operator removal
Suppression record While the objection remains in effect

The 30-day and 90-day periods from issue #84 apply only to successful and failed or cancelled DBOS workflow histories. That cleanup preserves domain evidence and raw artifacts. The raw-artifact periods above are the separate retention decision that issue #84 did not make.

Consequences

  • The directory stage remains in workflow 02 and can be ported from the PoC.
  • Existing adapters need follow-up changes to enforce this decision. Current differences do not block production use.
  • Browser fetching must stop treating a robots denial as blocked, minimize text and HTML before persistence, never store screenshots, and delete page artifacts after 30 days.
  • Directory extraction must distinguish company-context work contact points from private-only results and must not infer a person-to-phone link from text proximity alone.
  • Raw OpenRouter responses and unreadable response bodies must stop entering the raw store. (The zero-data-retention and no-training rule is account configuration, not a code gate — see the second 2026-08-19 amendment.)
  • DataForSEO payloads must be minimized before persistence and deleted after 30 days.
  • Work-phone reads and exports must include source and verification timestamps, and deletion requests require a suppression mechanism.
  • The raw bucket needs class-specific deletion because its existing retention and DBOS-history cleanup do not implement these periods.

Sources reviewed

Amendment 2026-08-19 — raw page and SERP artifacts are kept for 365 days (#329)

This amendment changes the retention period for minimized page and SERP artifacts in the raw store from 30 days to 365 days. Chris's decision on 2026-08-19: the project does not want to delete this data yet. Nothing else in the ADR changes — minimization before persistence, the prohibited-data rules of ADR-0006, transient screenshots, the never-stored raw LLM request and response, and the work-phone rules all stay as accepted on 2026-08-15.

Decision.

  1. Minimized first-party and directory page artifacts are retained for 365 days from retrieval, then deleted.
  2. Minimized DataForSEO SERP payloads are retained for 365 days from retrieval, then deleted.
  3. The deletion mechanism is unchanged and still owed: a class-specific lifecycle rule on the raw/dbos-web-* prefixes that also covers noncurrent versions, because the raw bucket has versioning enabled. Only the age condition differs from what the original decision implied.
  4. The period is provisional — "for now", in the words of the decision. It is a floor for re-extraction, not a claim that 365 days is the proportionate maximum. A later amendment may shorten it once the pipeline has run at #313 and #317 scale and the real volume is measured.

Reasoning.

  • 30 days contradicted the re-extraction promise. ADR-0014 states that "re-extraction can run from retained raw pages without a new scrape", and the extractor is versioned precisely so a new version can re-judge old evidence. With a 30-day window, every campaign older than a month is re-scrapeable only — a new extractor version would have to pay DataForSEO, the proxy pool, and the fetch again for pages the project already had.
  • The number now matches the minimized LLM parsed-result cache, which is already 365 days. One period for page, SERP, and LLM cache removes the case where the judgement outlives its evidence, which is the state that makes an audit record unreadable.
  • Minimization, not deletion, is the volume lever. Item 1 of #329 — full markdown and full HTML written per page, 90 KB of JSON for one live page — is where the bytes are. A 12-fold longer period on a minimized artifact in GCS is a cost the project can carry; a 12-fold longer period on the unminimized artifact is not. Minimization is therefore a harder prerequisite under this amendment than under the original 30-day rule, not a softer one.
  • The privacy position is unchanged in kind, longer in time. The retained artifact holds only company-context professional data after minimization, and the deletion and suppression rule still overrides the period: an objection or an operator removal deletes the artifact before 365 days. The longer period is accepted risk under the same legitimate-interest basis as the rest of this ADR, recorded here rather than assumed.

Revised retention summary (rows that change; all other rows in the table above stand):

Artifact or record Retention
Minimized first-party page text and HTML 365 days
Minimized directory page text and HTML 365 days
Minimized DataForSEO SERP payload 365 days

Consequences.

  • The consequences list above reads "delete page artifacts after 30 days" and "DataForSEO payloads must be minimized before persistence and deleted after 30 days". Read both as 365 days. The work itself — a lifecycle rule that covers noncurrent versions — is identical.
  • The immediate cleanup owed in #329 is not deferred by this amendment. The screenshot PNGs at raw/dbos-web-browser/screenshots/ have no retention period at all; ADR-0015 says a screenshot is never stored, so they are deleted on sight rather than aged out.
  • The raw LLM artifacts under raw/dbos-web-llm also gain no period from this amendment. They must stop being written and be removed, because the decision is that they are never stored.

Amendment 2026-08-19 — OpenRouter endpoint policy is account configuration, not a request-time gate (#329)

This amendment changes where the zero-data-retention and no-training rule is enforced, not whether it holds. Chris's decision on 2026-08-19: the providers allowed for a given model are configured on the OpenRouter account, and the code makes no provider-routing decision and does not fail closed on one.

Decision.

  1. The substantive rule stands: OpenRouter calls go to endpoints with zero data retention that do not train on prompts or responses, and OpenRouter prompt logging is disabled.
  2. The account is where that is set. Per-model provider selection and the account privacy settings are the control. The request carries the model and no provider-routing block.
  3. No request-time refusal. The transport does not classify a routing failure as a privacy refusal, and no code path fails an operation because a provider constraint could not be satisfied.
  4. The "fails closed when no compliant endpoint is available" sentence in the OpenRouter section above is superseded by this amendment.

Reasoning.

  • The control lives where the knowledge lives. The account knows which endpoints are permitted and stays right when a provider changes its policy. A copy of that decision in the request is stale the moment the account changes, and nothing tells the code it has gone stale.
  • A guard that blocks the work it guards is not protection. Enforced per request, this refused every LLM call in the first live tier m campaign — 25 of 30 companies — because the request-level constraint and the model's own first-party endpoint could not both be satisfied. No page text was protected by that; no enrichment happened either.
  • One layer, not three. Account configuration, a request-level constraint, and a bespoke terminal error class were three expressions of one rule, and the two in code could only disagree with the one that governs.

What does not change. Everything else in the OpenRouter section: full requests and raw provider responses are never stored, an unreadable response is not persisted as a diagnostic artifact, and the audit record keeps only the input-artifact hash, prompt and model versions, minimized parsed result, token use, cost, and evidence provenance. Those are about what is kept, and they are implemented (#329, PR #356).

Amendment 2026-08-28 — raw page artifacts keep full Markdown and HTML (#329)

This amendment withdraws the page-minimization requirement. Chris's decision on 2026-08-28: the project does not need to minimize fetched page artifacts, and the bucket storage cost of retaining the HTML is acceptable.

Decision.

  1. A browser-fetched first-party or directory page artifact keeps the full fetched Markdown and HTML.
  2. The browser adapter does not minimize the page body before it enters the raw store. This is a source-specific exception to ADR-0006.
  3. The artifact is still deleted 365 days from retrieval. The lifecycle rule must delete current objects and noncurrent versions from the versioned bucket.
  4. This amendment changes only browser-fetched page artifacts. Screenshots and raw OpenRouter requests and responses remain never-stored data. The separate DataForSEO rule is unchanged.

Reasoning.

  • Full Markdown and HTML preserve the fetched evidence for later extraction versions without another network request.
  • The existing artifact already stores both forms, and no new transformation layer is needed.
  • The expected bucket cost is acceptable. Storage volume is not a reason to remove HTML or add page-body minimization.

Consequences.

  • Item 1 of #329 is withdrawn. The current page writer already conforms to this amendment.
  • A retained artifact can contain personal data that the source published on the page. This knowingly accepts additional data-minimization and storage-limitation risk for the 365-day retention period.
  • An objection or operator deletion still overrides the age-based retention period.

Amendment 2026-08-28 — DataForSEO raw and structured storage are separate (#329)

This amendment corrects the earlier treatment of the 30-day SERP cache period as a retention period. Chris's decision on 2026-08-28: the complete raw DataForSEO response belongs in GCS, and its structured projection belongs in PostgreSQL for at least 365 days.

Decision.

  1. Each collected DataForSEO response is written unchanged to the raw GCS bucket as an immutable, content-addressed artifact. It is not reduced to the URL, title, and snippet fields before this write.
  2. PostgreSQL stores a queryable structured result set for each retrieval. It includes the query identity, provider task identity, retrieval time, ordered result fields, and the raw artifact URI and digest.
  3. Structured DataForSEO result sets are retained for at least 365 days. A later retrieval appends a new result set. It does not overwrite the prior result set.
  4. The 30-day SERP cache period is a freshness rule only. A result inside that period can satisfy a new request without another paid provider call. An older result causes a new retrieval, but it is not deleted because it is stale.
  5. The raw GCS artifact remains subject to the accepted 365-day raw-artifact lifecycle rule. An objection or operator deletion can remove data earlier.

Reasoning.

  • GCS preserves the provider evidence in its original form. PostgreSQL gives the application stable, queryable result history.
  • Freshness answers whether the application must acquire new evidence. Retention answers how long acquired evidence stays available. One value cannot safely implement both rules.
  • An immutable raw artifact and an append-only structured result set keep one retrieval traceable from a database row to the exact provider response.
  • The additional bucket storage cost is acceptable. Storage volume is not a reason to discard the raw provider response.

Consequences.

  • The current DataForSEO collection path does not conform. It writes only a compact result to web_content_cache and does not write the raw response to GCS.
  • web_content_cache can continue to select reusable results, but it cannot be the structured system of record. Its current upsert and expired-row purge must not delete DataForSEO result history before the 365-day minimum.
  • The earlier phrases "minimized DataForSEO SERP payload" and "SERP payloads are minimized before storage" are superseded. They do not apply to the raw GCS artifact. The structured PostgreSQL projection contains only the fields required by its read contract and provenance.