Skip to content

Externally-sourced people resolve through a mention-level alias layer, not the people table

Status: accepted Date: 2026-07-10 Amended: 2026-08-24 — preconditions met, implementation unblocked, and the resolution rule, contact storage, and suppression decided (issue #154) Deciders: Chris (solo founder) Depends on: ADR-0011 (local hub and ownership boundary) and ADR-0006 (minimization, registry-pure key rule) Unblocked by: ADR-0015, which supplies the lawful-basis decision decision 6 required Research: related-work synthesis, gap analysis

Context

ADR-0006 keys registry-sourced people on enhedsNummer and explicitly forbids solving external-people identity "by widening this rule". The future class-B source is specified (prephase SCRAPER/02: company-website team pages, own-gathered, vendors rejected) but is Phase-3 work — no scraper exists yet.

Deciding the identity contract now, while the CVR people model is being planned, means the future scraper builds against a stable target instead of improvising one — and means nothing in the CVR epics accidentally forecloses it. This ADR fixes the contract only; no tables ship until the scraper does.

The problem as scoped is company-scoped linkage, not population-wide deduplication: a scraped person arrives attached to one company (the site being crawled), so match candidates are that CVR's handful of registry people — a deterministic-first problem, not a probabilistic-clustering one.

Decision

  1. A person_mentions table, separate from people. One row per (source, company cvr, observed person): normalized name, observed title/role, per-field provenance (source, confidence, retrieved_at, evidence pointer to the raw snapshot), and a nullable resolved_person_id FK to people(person_id).
  2. people stays registry-pure. No synthetic person keys, no source discriminators in the hub's people PK space. An unresolved mention is a first-class standalone row, not a provisional person.
  3. Resolution is company-scoped and deterministic: the first given name and the surname must match exactly one of this CVR's registry people. Names are normalized by casefolding, punctuation removal, whitespace collapse, and folding æ ø å to ae oe aa. Zero matches or several matches resolve nothing. Probabilistic tooling is not adopted unless measured hit-rates demand it.

The accepted text used the SCRAPER/02 match key: normalized name plus role/title similarity, with published email or domain as tiebreaker. Role similarity is removed. A person can hold a board seat and be the CEO, so a role that disagrees is not evidence of a different human, and a fuzzy signal has no place in a rule that publishes an identity claim. The codebase already enforces unique-match-or-nothing for company mentions (enrichment/mentions.py) and person resolution now follows it.

Matching on first name and surname rather than the whole string is a deliberate relaxation, not a weakening. The register holds the full legal name (Jens Peter Hansen) and a website almost always drops the middle name (Jens Hansen), so a whole-string match would fail on the common case. Uniqueness inside one CVR's handful of people is what carries the decision: a company holding both a Jens Hansen and a Jens Peter Hansen produces two matches, so nothing resolves. 4. Conflict = keep both, ranked. Registry-sourced fields outrank scraped fields; a wrong resolution is reversed by clearing resolved_person_id — merges are auditable and reversible, never destructive. 5. No cross-company person identity. The same human at two companies is two mentions. A person-level cluster ID across companies (or across sources without a registry key) is a new ADR with its own lawful-basis analysis — the minimization posture argues against building it speculatively. 6. Preconditions before any implementation stores a mention. Both are now met, and implementation is unblocked: - ADR-0006 minimization applies at the scraper's adapter boundary (role-relevant fields only; no scraping of residential or private data). - A recorded lawful-basis decision for scraped people data — legitimate interest analysis plus the GDPR Art. 14 information duty and subject-rights handling (prephase SCRAPER/02 legal section). Storage ships after that decision, not before.

ADR-0015 supplies the second one. It records legitimate interest as the basis for first-party professional person data, and it decides the Article 14 posture explicitly: the controller does not proactively notify, and accepts that as an unresolved compliance risk rather than a finding that an exception applies. That is a recorded decision, which is what this precondition asked for.

  1. A person's work contact points are stored separately from the mention, one row per value. A phone or an email attaches to a person only when the source states the attribution — through text, page structure, structured data, or an address whose own name matches the mention. Without that statement the value is a company work contact point and the person stays a bare mention. This is ADR-0015's "no inference from proximity alone" rule, and it now covers email as well as phone.

One row per value rather than columns on the mention, for two reasons. Decision 8's objection deletes contact rows and leaves the mention and its evidence intact, which is exactly what ADR-0015 asks for. And a new contact kind then needs no migration. It is also the shape web_company_contact already uses, so the two contact tables read the same way.

  1. An objection deletes the person's work contact data and writes a suppression record that holds the objected value in clear. Collection consults it before it stores a value, so a later refresh cannot re-add what an objection removed.

Storing the value in clear is a knowing trade. A hash was the alternative, and a plain hash of an eight-digit Danish phone number brute-forces in seconds, so only an HMAC with a secret pepper would have protected it — one more secret, and one more thing whose loss makes every suppression unenforceable. The consequence of the choice is that the suppression table is itself personal data: it is a readable list of the people who objected, and it carries the same row-level security as the rest of the hub.

No inbound objection path and no operator command ship with the first implementation. An objection is handled by hand, from a runbook statement that performs the deletion and the suppression insert in one transaction, so an operator cannot do half of it. This is a deliberate deferral on the grounds that no objection has been received; the risk accepted is that a hand-written deletion is slower than a tool.

Options considered

  • Mention-level table + optional FK (chosen). Registry PK space stays pure; unresolved data is visible as unresolved; reversible merges.
  • Alias rows inside people (rejected). Widens the enhedsNummer key space — exactly what ADR-0006 forbids; every consumer would need to re-learn what a "person" row means.
  • Full probabilistic ER with clustering now (rejected). Solves a problem no committed scenario has, on data that does not exist, with the hardest GDPR posture.

Consequences

  • The CVR epics need no schema change for this — the contract is additive (person_mentions arrives with the scraper). What the CVR work must preserve is only people's registry-pure key rule, already in ADR-0006.
  • The future scraper's output signal shape (SCRAPER/02's person_role/person_contact sketch) must reference mentions, with provenance ∈ {registry, discovered} edges kept distinct.
  • Downstream consumers can trust: a people row is always a registry entity; a mention without resolved_person_id is observed-but-unlinked; nothing is silently merged.
  • The unresolved mention is the normal case, not the exception. This ADR was written expecting resolution against registry people. A salgschef, a projektleder, and a receptionist hold no CVR role at all, so most people a website publishes will never resolve. Decision 2 therefore carries the weight of the design: unresolved mentions are the bulk of the output and must be readable, not a waiting room for resolution. company_detail gains a people section that lists every mention with its title, contacts, source URL, collection time, and its registry person where one was matched.
  • A resolved mention publishes no relationship edge. The registry role is already the edge, and restating it as an observed edge would assert the same fact twice from two source classes. Resolution sets resolved_person_id and nothing more.

Resolved questions

  • Where person_mentions lives: this repository. Web enrichment moved here with ADR-0016, so the shortlist-scoped argument for the enrichment repository no longer points anywhere else. The hub-is-sole-writer rule from ADR-0011 decides it.
  • Retention for stale mentions: no age-based expiry. ADR-0015 already decided this for the work phone it governs — a retained work contact point remains available indefinitely unless an objection or another explicit deletion rule applies, and later removal from the source does not delete or demote it. The mention follows its contacts. SCRAPER/02's proposed 90-day re-crawl TTL is not adopted; freshness is a refresh decision, and every displayed value already carries its source URL and collection time so a consumer can see its age.