Externally-sourced people resolve through a mention-level alias layer, not the people table¶
Status: accepted Date: 2026-07-10 Amended: 2026-08-24 — preconditions met, implementation unblocked, and the resolution rule, contact storage, and suppression decided (issue #154) Deciders: Chris (solo founder) Depends on: ADR-0011 (local hub and ownership boundary) and ADR-0006 (minimization, registry-pure key rule) Unblocked by: ADR-0015, which supplies the lawful-basis decision decision 6 required Research: related-work synthesis, gap analysis
Context¶
ADR-0006 keys registry-sourced people on enhedsNummer and explicitly
forbids solving external-people identity "by widening this rule". The future
class-B source is specified (prephase SCRAPER/02: company-website team pages,
own-gathered, vendors rejected) but is Phase-3 work — no scraper exists yet.
Deciding the identity contract now, while the CVR people model is being planned, means the future scraper builds against a stable target instead of improvising one — and means nothing in the CVR epics accidentally forecloses it. This ADR fixes the contract only; no tables ship until the scraper does.
The problem as scoped is company-scoped linkage, not population-wide deduplication: a scraped person arrives attached to one company (the site being crawled), so match candidates are that CVR's handful of registry people — a deterministic-first problem, not a probabilistic-clustering one.
Decision¶
- A
person_mentionstable, separate frompeople. One row per (source, companycvr, observed person): normalized name, observed title/role, per-field provenance (source,confidence,retrieved_at, evidence pointer to the raw snapshot), and a nullableresolved_person_idFK topeople(person_id). peoplestays registry-pure. No synthetic person keys, no source discriminators in the hub's people PK space. An unresolved mention is a first-class standalone row, not a provisional person.- Resolution is company-scoped and deterministic: the first given name and
the surname must match exactly one of this CVR's registry people. Names
are normalized by casefolding, punctuation removal, whitespace collapse,
and folding
æ ø åtoae oe aa. Zero matches or several matches resolve nothing. Probabilistic tooling is not adopted unless measured hit-rates demand it.
The accepted text used the SCRAPER/02 match key: normalized name plus
role/title similarity, with published email or domain as tiebreaker. Role
similarity is removed. A person can hold a board seat and be the CEO, so a
role that disagrees is not evidence of a different human, and a fuzzy
signal has no place in a rule that publishes an identity claim. The
codebase already enforces unique-match-or-nothing for company mentions
(enrichment/mentions.py) and person resolution now follows it.
Matching on first name and surname rather than the whole string is a
deliberate relaxation, not a weakening. The register holds the full legal
name (Jens Peter Hansen) and a website almost always drops the middle
name (Jens Hansen), so a whole-string match would fail on the common
case. Uniqueness inside one CVR's handful of people is what carries the
decision: a company holding both a Jens Hansen and a Jens Peter Hansen
produces two matches, so nothing resolves.
4. Conflict = keep both, ranked. Registry-sourced fields outrank scraped
fields; a wrong resolution is reversed by clearing resolved_person_id —
merges are auditable and reversible, never destructive.
5. No cross-company person identity. The same human at two companies is
two mentions. A person-level cluster ID across companies (or across
sources without a registry key) is a new ADR with its own lawful-basis
analysis — the minimization posture argues against building it
speculatively.
6. Preconditions before any implementation stores a mention. Both are now
met, and implementation is unblocked:
- ADR-0006 minimization applies at the scraper's adapter boundary
(role-relevant fields only; no scraping of residential or private data).
- A recorded lawful-basis decision for scraped people data — legitimate
interest analysis plus the GDPR Art. 14 information duty and
subject-rights handling (prephase SCRAPER/02 legal section). Storage
ships after that decision, not before.
ADR-0015 supplies the second one. It records legitimate interest as the basis for first-party professional person data, and it decides the Article 14 posture explicitly: the controller does not proactively notify, and accepts that as an unresolved compliance risk rather than a finding that an exception applies. That is a recorded decision, which is what this precondition asked for.
- A person's work contact points are stored separately from the mention, one row per value. A phone or an email attaches to a person only when the source states the attribution — through text, page structure, structured data, or an address whose own name matches the mention. Without that statement the value is a company work contact point and the person stays a bare mention. This is ADR-0015's "no inference from proximity alone" rule, and it now covers email as well as phone.
One row per value rather than columns on the mention, for two reasons.
Decision 8's objection deletes contact rows and leaves the mention and its
evidence intact, which is exactly what ADR-0015 asks for. And a new contact
kind then needs no migration. It is also the shape web_company_contact
already uses, so the two contact tables read the same way.
- An objection deletes the person's work contact data and writes a suppression record that holds the objected value in clear. Collection consults it before it stores a value, so a later refresh cannot re-add what an objection removed.
Storing the value in clear is a knowing trade. A hash was the alternative, and a plain hash of an eight-digit Danish phone number brute-forces in seconds, so only an HMAC with a secret pepper would have protected it — one more secret, and one more thing whose loss makes every suppression unenforceable. The consequence of the choice is that the suppression table is itself personal data: it is a readable list of the people who objected, and it carries the same row-level security as the rest of the hub.
No inbound objection path and no operator command ship with the first implementation. An objection is handled by hand, from a runbook statement that performs the deletion and the suppression insert in one transaction, so an operator cannot do half of it. This is a deliberate deferral on the grounds that no objection has been received; the risk accepted is that a hand-written deletion is slower than a tool.
Options considered¶
- Mention-level table + optional FK (chosen). Registry PK space stays pure; unresolved data is visible as unresolved; reversible merges.
- Alias rows inside
people(rejected). Widens theenhedsNummerkey space — exactly what ADR-0006 forbids; every consumer would need to re-learn what a "person" row means. - Full probabilistic ER with clustering now (rejected). Solves a problem no committed scenario has, on data that does not exist, with the hardest GDPR posture.
Consequences¶
- The CVR epics need no schema change for this — the contract is additive
(
person_mentionsarrives with the scraper). What the CVR work must preserve is onlypeople's registry-pure key rule, already in ADR-0006. - The future scraper's output signal shape (SCRAPER/02's
person_role/person_contactsketch) must reference mentions, withprovenance ∈ {registry, discovered}edges kept distinct. - Downstream consumers can trust: a
peoplerow is always a registry entity; a mention withoutresolved_person_idis observed-but-unlinked; nothing is silently merged. - The unresolved mention is the normal case, not the exception. This ADR
was written expecting resolution against registry people. A salgschef, a
projektleder, and a receptionist hold no CVR role at all, so most people a
website publishes will never resolve. Decision 2 therefore carries the
weight of the design: unresolved mentions are the bulk of the output and
must be readable, not a waiting room for resolution.
company_detailgains a people section that lists every mention with its title, contacts, source URL, collection time, and its registry person where one was matched. - A resolved mention publishes no relationship edge. The registry role is
already the edge, and restating it as an
observededge would assert the same fact twice from two source classes. Resolution setsresolved_person_idand nothing more.
Resolved questions¶
- Where
person_mentionslives: this repository. Web enrichment moved here with ADR-0016, so the shortlist-scoped argument for the enrichment repository no longer points anywhere else. The hub-is-sole-writer rule from ADR-0011 decides it. - Retention for stale mentions: no age-based expiry. ADR-0015 already decided this for the work phone it governs — a retained work contact point remains available indefinitely unless an objection or another explicit deletion rule applies, and later removal from the source does not delete or demote it. The mention follows its contacts. SCRAPER/02's proposed 90-day re-crawl TTL is not adopted; freshness is a refresh decision, and every displayed value already carries its source URL and collection time so a consumer can see its age.