Skip to content

Next expansion — related work synthesis

Date: 2026-07-10 Question this supports: what the next planning cycle after the CVR ingester (PRD #15, issues #16–#26) should commit to, across the three named follow-ups: the Regnskabsdata ingester, the CVR_Events sub-day freshness upgrade, and the entity-resolution/alias layer for externally-sourced people. Companion: gap analysis

Evidence classes used below: [live] = probed today against the real endpoint; [repo] = recorded in this platform's docs (captured 2026-06, live-tested then); [prephase] = prephase mapping/ADR content (analysis, not live-tested); [inference] = our synthesis, marked as such.

The three candidates are not the same kind of work

Regnskabsdata CVR_Events People entity-resolution
Kind of work Second ingester (build) Freshness upgrade to an existing ingester (build + new upstream) Cross-cutting identity layer (design + build)
Documented gate None — named "next" in docs/wip/cvr-ingester-plan.md ADR-0005: deferred upgrade path, "gated on evidence"; trigger metric an open question ADR-0006: deferred "until a class-B source lands"; must not widen the enhedsNummer key rule
Gate status today Open Closed — no product requirement for sub-day freshness exists yet Closed — no class-B people source is ingested or scheduled before Phase 3
Upstream access Free, no credential [live] Datafordeler API key held; person entities still OAuth+MitID-gated [repo] n/a (internal layer; consumes future scraper output)
Prephase maturity Full schema mapping (lassox doc 04) [prephase] Surface identified, event schema unexplored [repo] Design sketch exists (SCRAPER/02) [prephase]
Depends on shipped code core/ framework (E1) E1–E3 + a bulk baseline E3 (people/people_roles) + a scraper that doesn't exist in this repo

The synthesis conclusion up front: only Regnskabsdata is build-shaped today. The other two are decision-shaped — what they need from this cycle is an explicit, recorded trigger condition, not an implementation plan. Planning all three to the same depth would contradict two accepted ADRs.

Theme 1 — Regnskabsdata: the framework-reuse test case

What is established

  • Surface [live, probed 2026-07-10]: http://distribution.virk.dk/offentliggoerelser/_search answers HTTP 200 with no credential. The alias resolves to index indberetninger-20230623 (also aliased indberetninger, offentliggoerelser-prod). Total: 6,410,977 publication records.
  • Document shape [live] matches prephase doc 04: cvrNummer, sagsNummer, regnskab.regnskabsperiode.{startDato,slutDato}, sidstOpdateret, indlaesningsTidspunkt, offentliggoerelsesTidspunkt, omgoerelse, offentliggoerelsestype, and nested dokumenter[] (dokumentType e.g. AARSRAPPORT, dokumentMimeType, dokumentUrl) — typically one PDF and one XBRL (application/xml) URL per report, served from http://regnskaber.virk.dk/….
  • The cluster is Elasticsearch 6.8.23 [live]: the unauthenticated root of distribution.virk.dk reports version.number: 6.8.23 (Erst.Distribution.AWS.Prod.Cluster, AWS eu-north-1). This contradicts the "behaves like ES 1.7" claim in docs/data-sources/cvr-elasticsearch.md and the "ES 1.7 dialect is a contract" consequence in ADR-0005 — both sourced from Erhvervsstyrelsen's getting-started guide, which is evidently outdated. Consequence if confirmed against cvr-permanent with auth: search_after, sliced scroll, and modern query DSL are available to both ingesters. Logged as a gap (G1) rather than an ADR edit because the authenticated index behavior is what counts.
  • Two-tier dataflow [prephase, doc 04]: metadata is bulk (upstream updates ~10 min; re-scroll on sidstOpdateret), documents are heavy and lazy (fetch XBRL/PDF on demand, parse key figures, stamp details_cached_at). This is the prephase ADR-0004/0005 split applied verbatim.
  • Hub target [prephase, doc 04]: financial_reports (metadata, 1:N to companies) + financial_metrics (promoted nøgletal) + raw XBRL facts as JSONB, lazy-filled.
  • Reports are effectively immutable once published; corrections arrive as new publications flagged omgoerelse [prephase]. Long detail TTL justified.

What the framework already gives it vs. what is genuinely new

Reuse (built in E0–E4 of the CVR plan): core/config, core/db, core/raw_store (run-scoped NDJSON.gz to GCS), core/ledger (ingestion_runs + watermark), staging+swap and upsert loaders, Cloud Run Job + Scheduler runtime, Supabase migration path. The adapter/ingester split holds: a new adapters/regnskab_es + ingesters/regnskab package, core/ unchanged — this is the stated purpose of the framework (docs/wip/cvr-ingester-plan.md, Scope).

New capabilities Regnskabsdata forces, in rising order of novelty [inference]:

  1. No-credential adapter variant — simpler than cvr_es, but the HTTP-only constraint (and its "run only from Cloud Run" consequence) still applies: regnskaber.virk.dk document URLs are plain HTTP too [live].
  2. A second watermark discipline on an index where sidstOpdateret is the cursor but indlaesningsId — the prephase idempotency key — is missing on 3,440,462 of 6,410,977 records (54%) [live]. The ES _id (urn:ofk:oid:…, present on every hit) is the durable key; prephase doc 04's key choice is superseded.
  3. The first lazy-detail path in the platform: on-demand document fetch + parse + details_cached_at stamping. The CVR ingester is bulk-only, so core/ has no machinery for this yet — this is the main framework extension, and it is exactly the pattern every later detail-heavy source (Tinglysning, EMO, …) will reuse.
  4. XBRL key-figure extraction — the only genuinely new skill in the slice. Danish filings use Erhvervsstyrelsen's yearly-versioned XBRL taxonomies (DK GAAP and IFRS variants), so the mapping XBRL → financial_metrics is a versioned contract needing a per-taxonomy extractor [prephase, doc 04 open question]. Candidate tooling: Arelle (the de-facto open-source XBRL processor) vs. a narrow hand-rolled element extractor for the ~10 promoted nøgletal [inference — needs a spike against real DK filings; G4].

Contradictions / tensions surfaced

  • The ES-version contradiction above (guide says 1.7, cluster says 6.8.23).
  • Prephase doc 04 proposes unique (cvr, period_start, period_end) and append-style correction handling — these conflict; omgoerelse handling must pick one (G3).
  • ~~The index may carry non-regnskab publication types~~ — resolved [live]: all 6,410,977 records are offentliggoerelsestype: regnskab; no population filter needed. XBRL coverage: 4,028,669 records (63%) carry an XML document; the rest are PDF/paper-only, bounding financial_metrics coverage (~37% of reports will have null metrics, mostly older filings).

Theme 2 — CVR_Events: an upgrade with no trigger, and no forcing function

What is established

  • ADR-0005 already chose the upgrade shape when it becomes justified: CVR_Events as trigger + ES re-fetch of the full document — not events + GraphQL as the data path (rejected: second full mapping, person entities OAuth-gated) [repo, ADR-0005 options table].
  • The upstream is Datafordeleren's entity-based events model (introduced late 2025, ~30s cadence, "vedligeholdelse af kopiregister med hændelser"); legacy Datafordeler CVR channels deprecate 2027-01-15 [repo, ADR-0005 / cvr-datafordeler.md].
  • The 2027 deprecation is not a forcing function for us [inference]: our v1 surface is distribution.virk.dk, not Datafordeleren. Nothing we run breaks in 2027. The deprecation only matters if/when we adopt events — it tells us which events surface to adopt (the entity-based model), a question ADR-0005 already answered.
  • The event schema, subscription/cursor mechanics, and delivery guarantees are unexplored — no live test exists in docs/data-sources/ [repo, absence].

What the literature/pattern says

Batch-polling → event-notification for mirror maintenance is the classic derived-data upgrade (DDIA's "keeping systems in sync"): events reduce detection latency but add subscription state, gap/replay handling, and a second failure domain — and the monthly reconciliation scroll remains necessary as the correctness backstop either way [inference from platform's own ADR framing]. The cost is therefore permanent (two upstreams, two credentials, cursor state) while the benefit is bounded by an actual product need for sub-day detection. ADR-0005 names ownership-change monitoring as the likely trigger.

What this cycle should produce for CVR_Events [inference]: not a build plan but a recorded trigger contract — e.g. "adopt when a paying product scenario requires change detection < X hours, or when daily delta measurably misses/lags by more than Y" — plus, optionally, a cheap evidence spike: measure the actual lag of the daily sidstOpdateret delta once E1/E3 run in test (the ledger gives this for free). Decision inputs, not code.

Theme 3 — People entity-resolution: design exists, its input doesn't

What is established

  • ADR-0006 keys registry people on enhedsNummer and explicitly forbids solving external-people identity "by widening this rule" — an alias/ resolution layer is the named mechanism, deferred until a class-B source lands [repo, ADR-0006].
  • The class-B source is specified in prephase SCRAPER/02 (own-gathered only: L0 registry + L1 company-website team pages; vendors rejected on quality) and is Phase 3 work — no scraper exists in this repository, and the scraper pipeline itself belongs to enrichment, not the CVR hub epics [prephase].
  • SCRAPER/02 already sketches the resolution design [prephase]:
  • match key: normalized name + company (CVR) + role/title similarity; published email/domain as tiebreaker;
  • per-field provenance (source, confidence, retrieved_at), best-per-field merge, registry fields outrank scraped;
  • conflict = keep both, ranked — never silently overwrite; wrong merges stay auditable and reversible;
  • registry edges vs. discovered edges carry provenance per prephase ADR-0006 (graph strategy).
  • Legal precondition [prephase SCRAPER/02]: storing scraped people data makes us a GDPR controller — lawful basis + Art. 14 information duty + subject rights must be decided before storage, extending ADR-0006's minimization rule to a source where the data is not registry-published.

What the pattern says

The problem as scoped is company-scoped linkage, not population-wide dedup [inference]: candidates for a scraped person are the handful of registry people already attached to that CVR, so deterministic/rule-based matching with the SCRAPER/02 key likely dominates; probabilistic linkage tooling (e.g. Splink) only earns its complexity if cross-company person identity ("same human, two companies, no registry key") becomes a product need — which the minimization posture actively discourages. The schema consequence is additive and small: an people_aliases (or person_mentions) table pointing at people.person_id where resolved, standing alone where not — the hub's people PK space stays registry-pure, exactly as ADR-0006 requires.

What this cycle should produce for ER [inference]: at most a short design ADR fixing the alias-table shape and the "keep both, ranked" conflict rule as a contract for the future scraper, plus the lawful-basis homework it requires — so that Phase 3 can build against a stable identity model. Building the layer now would ship code with no caller and no data.

  1. Regnskabsdata: full flow (grill → PRD → issues). It is the framework's first reuse test, its gate is open, its access is free, and its two genuinely new pieces (lazy-detail path, XBRL extractor) are exactly the capabilities later sources need. Sequencing constraint: it consumes core/ as built by E1 and references companies FKs — it can be planned now but implemented only after E1 lands (serial execution per ADR-0003).
  2. CVR_Events: no PRD. Record the trigger contract (amend ADR-0005's open question into a measurable condition) and add a zero-cost lag measurement to the ops epic already planned (E4/#25 territory).
  3. Entity-resolution: no PRD. One design ADR (alias layer + provenance + conflict rule + lawful-basis precondition) so the identity contract exists before any scraper work starts.

This keeps the cycle honest against ADR-0005/0006's own gates while still moving all three forward in the way each actually needs.

Handoff

Next step per the feature flow: /grill-with-docs on the scope above — the one genuinely open decision is whether Chris accepts the "one PRD + two decision records" shape or wants full plans for all three. Companion gap list: 2026-07-10-next-expansion-gaps.md.