Next expansion — related work synthesis¶
Date: 2026-07-10
Question this supports: what the next planning cycle after the CVR ingester
(PRD #15, issues #16–#26) should commit to, across the three named follow-ups:
the Regnskabsdata ingester, the CVR_Events sub-day freshness upgrade,
and the entity-resolution/alias layer for externally-sourced people.
Companion: gap analysis
Evidence classes used below: [live] = probed today against the real endpoint; [repo] = recorded in this platform's docs (captured 2026-06, live-tested then); [prephase] = prephase mapping/ADR content (analysis, not live-tested); [inference] = our synthesis, marked as such.
The three candidates are not the same kind of work¶
| Regnskabsdata | CVR_Events | People entity-resolution | |
|---|---|---|---|
| Kind of work | Second ingester (build) | Freshness upgrade to an existing ingester (build + new upstream) | Cross-cutting identity layer (design + build) |
| Documented gate | None — named "next" in docs/wip/cvr-ingester-plan.md |
ADR-0005: deferred upgrade path, "gated on evidence"; trigger metric an open question | ADR-0006: deferred "until a class-B source lands"; must not widen the enhedsNummer key rule |
| Gate status today | Open | Closed — no product requirement for sub-day freshness exists yet | Closed — no class-B people source is ingested or scheduled before Phase 3 |
| Upstream access | Free, no credential [live] | Datafordeler API key held; person entities still OAuth+MitID-gated [repo] | n/a (internal layer; consumes future scraper output) |
| Prephase maturity | Full schema mapping (lassox doc 04) [prephase] | Surface identified, event schema unexplored [repo] | Design sketch exists (SCRAPER/02) [prephase] |
| Depends on shipped code | core/ framework (E1) |
E1–E3 + a bulk baseline | E3 (people/people_roles) + a scraper that doesn't exist in this repo |
The synthesis conclusion up front: only Regnskabsdata is build-shaped today. The other two are decision-shaped — what they need from this cycle is an explicit, recorded trigger condition, not an implementation plan. Planning all three to the same depth would contradict two accepted ADRs.
Theme 1 — Regnskabsdata: the framework-reuse test case¶
What is established¶
- Surface [live, probed 2026-07-10]:
http://distribution.virk.dk/offentliggoerelser/_searchanswers HTTP 200 with no credential. The alias resolves to indexindberetninger-20230623(also aliasedindberetninger,offentliggoerelser-prod). Total: 6,410,977 publication records. - Document shape [live] matches prephase doc 04:
cvrNummer,sagsNummer,regnskab.regnskabsperiode.{startDato,slutDato},sidstOpdateret,indlaesningsTidspunkt,offentliggoerelsesTidspunkt,omgoerelse,offentliggoerelsestype, and nesteddokumenter[](dokumentTypee.g.AARSRAPPORT,dokumentMimeType,dokumentUrl) — typically one PDF and one XBRL (application/xml) URL per report, served fromhttp://regnskaber.virk.dk/…. - The cluster is Elasticsearch 6.8.23 [live]: the unauthenticated root of
distribution.virk.dkreportsversion.number: 6.8.23(Erst.Distribution.AWS.Prod.Cluster, AWS eu-north-1). This contradicts the "behaves like ES 1.7" claim indocs/data-sources/cvr-elasticsearch.mdand the "ES 1.7 dialect is a contract" consequence in ADR-0005 — both sourced from Erhvervsstyrelsen's getting-started guide, which is evidently outdated. Consequence if confirmed againstcvr-permanentwith auth:search_after, sliced scroll, and modern query DSL are available to both ingesters. Logged as a gap (G1) rather than an ADR edit because the authenticated index behavior is what counts. - Two-tier dataflow [prephase, doc 04]: metadata is bulk (upstream updates
~10 min; re-scroll on
sidstOpdateret), documents are heavy and lazy (fetch XBRL/PDF on demand, parse key figures, stampdetails_cached_at). This is the prephase ADR-0004/0005 split applied verbatim. - Hub target [prephase, doc 04]:
financial_reports(metadata, 1:N tocompanies) +financial_metrics(promoted nøgletal) + raw XBRL facts as JSONB, lazy-filled. - Reports are effectively immutable once published; corrections arrive as
new publications flagged
omgoerelse[prephase]. Long detail TTL justified.
What the framework already gives it vs. what is genuinely new¶
Reuse (built in E0–E4 of the CVR plan): core/config, core/db,
core/raw_store (run-scoped NDJSON.gz to GCS), core/ledger
(ingestion_runs + watermark), staging+swap and upsert loaders, Cloud Run
Job + Scheduler runtime, Supabase migration path. The adapter/ingester split
holds: a new adapters/regnskab_es + ingesters/regnskab package,
core/ unchanged — this is the stated purpose of the framework
(docs/wip/cvr-ingester-plan.md, Scope).
New capabilities Regnskabsdata forces, in rising order of novelty [inference]:
- No-credential adapter variant — simpler than
cvr_es, but the HTTP-only constraint (and its "run only from Cloud Run" consequence) still applies:regnskaber.virk.dkdocument URLs are plain HTTP too [live]. - A second watermark discipline on an index where
sidstOpdateretis the cursor butindlaesningsId— the prephase idempotency key — is missing on 3,440,462 of 6,410,977 records (54%) [live]. The ES_id(urn:ofk:oid:…, present on every hit) is the durable key; prephase doc 04's key choice is superseded. - The first lazy-detail path in the platform: on-demand document fetch +
parse +
details_cached_atstamping. The CVR ingester is bulk-only, socore/has no machinery for this yet — this is the main framework extension, and it is exactly the pattern every later detail-heavy source (Tinglysning, EMO, …) will reuse. - XBRL key-figure extraction — the only genuinely new skill in the
slice. Danish filings use Erhvervsstyrelsen's yearly-versioned XBRL
taxonomies (DK GAAP and IFRS variants), so the mapping XBRL →
financial_metricsis a versioned contract needing a per-taxonomy extractor [prephase, doc 04 open question]. Candidate tooling: Arelle (the de-facto open-source XBRL processor) vs. a narrow hand-rolled element extractor for the ~10 promoted nøgletal [inference — needs a spike against real DK filings; G4].
Contradictions / tensions surfaced¶
- The ES-version contradiction above (guide says 1.7, cluster says 6.8.23).
- Prephase doc 04 proposes
unique (cvr, period_start, period_end)and append-style correction handling — these conflict;omgoerelsehandling must pick one (G3). - ~~The index may carry non-regnskab publication types~~ — resolved [live]:
all 6,410,977 records are
offentliggoerelsestype: regnskab; no population filter needed. XBRL coverage: 4,028,669 records (63%) carry an XML document; the rest are PDF/paper-only, boundingfinancial_metricscoverage (~37% of reports will have null metrics, mostly older filings).
Theme 2 — CVR_Events: an upgrade with no trigger, and no forcing function¶
What is established¶
- ADR-0005 already chose the upgrade shape when it becomes justified:
CVR_Eventsas trigger + ES re-fetch of the full document — not events + GraphQL as the data path (rejected: second full mapping, person entities OAuth-gated) [repo, ADR-0005 options table]. - The upstream is Datafordeleren's entity-based events model (introduced late 2025, ~30s cadence, "vedligeholdelse af kopiregister med hændelser"); legacy Datafordeler CVR channels deprecate 2027-01-15 [repo, ADR-0005 / cvr-datafordeler.md].
- The 2027 deprecation is not a forcing function for us [inference]: our
v1 surface is
distribution.virk.dk, not Datafordeleren. Nothing we run breaks in 2027. The deprecation only matters if/when we adopt events — it tells us which events surface to adopt (the entity-based model), a question ADR-0005 already answered. - The event schema, subscription/cursor mechanics, and delivery guarantees
are unexplored — no live test exists in
docs/data-sources/[repo, absence].
What the literature/pattern says¶
Batch-polling → event-notification for mirror maintenance is the classic derived-data upgrade (DDIA's "keeping systems in sync"): events reduce detection latency but add subscription state, gap/replay handling, and a second failure domain — and the monthly reconciliation scroll remains necessary as the correctness backstop either way [inference from platform's own ADR framing]. The cost is therefore permanent (two upstreams, two credentials, cursor state) while the benefit is bounded by an actual product need for sub-day detection. ADR-0005 names ownership-change monitoring as the likely trigger.
What this cycle should produce for CVR_Events [inference]: not a build
plan but a recorded trigger contract — e.g. "adopt when a paying product
scenario requires change detection < X hours, or when daily delta measurably
misses/lags by more than Y" — plus, optionally, a cheap evidence spike:
measure the actual lag of the daily sidstOpdateret delta once E1/E3 run in
test (the ledger gives this for free). Decision inputs, not code.
Theme 3 — People entity-resolution: design exists, its input doesn't¶
What is established¶
- ADR-0006 keys registry people on
enhedsNummerand explicitly forbids solving external-people identity "by widening this rule" — an alias/ resolution layer is the named mechanism, deferred until a class-B source lands [repo, ADR-0006]. - The class-B source is specified in prephase SCRAPER/02 (own-gathered only: L0 registry + L1 company-website team pages; vendors rejected on quality) and is Phase 3 work — no scraper exists in this repository, and the scraper pipeline itself belongs to enrichment, not the CVR hub epics [prephase].
- SCRAPER/02 already sketches the resolution design [prephase]:
- match key: normalized name + company (CVR) + role/title similarity; published email/domain as tiebreaker;
- per-field provenance (source, confidence,
retrieved_at), best-per-field merge, registry fields outrank scraped; - conflict = keep both, ranked — never silently overwrite; wrong merges stay auditable and reversible;
- registry edges vs. discovered edges carry
provenanceper prephase ADR-0006 (graph strategy). - Legal precondition [prephase SCRAPER/02]: storing scraped people data makes us a GDPR controller — lawful basis + Art. 14 information duty + subject rights must be decided before storage, extending ADR-0006's minimization rule to a source where the data is not registry-published.
What the pattern says¶
The problem as scoped is company-scoped linkage, not population-wide
dedup [inference]: candidates for a scraped person are the handful of
registry people already attached to that CVR, so deterministic/rule-based
matching with the SCRAPER/02 key likely dominates; probabilistic linkage
tooling (e.g. Splink) only earns its complexity if cross-company person
identity ("same human, two companies, no registry key") becomes a product
need — which the minimization posture actively discourages. The schema
consequence is additive and small: an people_aliases (or
person_mentions) table pointing at people.person_id where resolved,
standing alone where not — the hub's people PK space stays registry-pure,
exactly as ADR-0006 requires.
What this cycle should produce for ER [inference]: at most a short design ADR fixing the alias-table shape and the "keep both, ranked" conflict rule as a contract for the future scraper, plus the lawful-basis homework it requires — so that Phase 3 can build against a stable identity model. Building the layer now would ship code with no caller and no data.
Synthesis — recommended shape of the next cycle¶
- Regnskabsdata: full flow (grill → PRD → issues). It is the framework's
first reuse test, its gate is open, its access is free, and its two
genuinely new pieces (lazy-detail path, XBRL extractor) are exactly the
capabilities later sources need. Sequencing constraint: it consumes
core/as built by E1 and referencescompaniesFKs — it can be planned now but implemented only after E1 lands (serial execution per ADR-0003). - CVR_Events: no PRD. Record the trigger contract (amend ADR-0005's open question into a measurable condition) and add a zero-cost lag measurement to the ops epic already planned (E4/#25 territory).
- Entity-resolution: no PRD. One design ADR (alias layer + provenance + conflict rule + lawful-basis precondition) so the identity contract exists before any scraper work starts.
This keeps the cycle honest against ADR-0005/0006's own gates while still moving all three forward in the way each actually needs.
Handoff¶
Next step per the feature flow: /grill-with-docs on the scope above —
the one genuinely open decision is whether Chris accepts the
"one PRD + two decision records" shape or wants full plans for all three.
Companion gap list: 2026-07-10-next-expansion-gaps.md.