CVR ingestion: single Elasticsearch surface, incremental refresh, near-source storage¶
Status: accepted
Date: 2026-07-10
Deciders: Chris (solo founder)
Depends on: ADR-0011 (local hub and ownership boundary)
Supersedes: the prephase DATAFORDELER/README.md guidance that Datafordeleren is the main source for bulk CVR acquisition
Current storage decision: promoted columns and raw_snapshot_uri in
PostgreSQL; source documents in GCS. This replaces decision 3's original raw
JSONB choice, as recorded in the 2026-08-12 amendments below.
Context¶
The CVR ingester is the first ingester built in this repository and feeds the
anchor of the hub: companies, production_units, people, and
people_roles. Two independent upstream surfaces exist, both live-tested
(docs/data-sources/ in the platform repo):
distribution.virk.dk— Erhvervsstyrelsen's Elasticsearch 1.7 index (cvr-permanent), one free Basic Auth credential, three document types (virksomhed~2.26M,produktionsenhed~2.87M,deltager~1.82M), full bulk extraction via the scroll API, delta queries viasidstOpdateret. HTTP only — the endpoint is not reachable over HTTPS, so credentials travel in cleartext.- Datafordeleren — CVR GraphQL v2 plus the entity-based
CVR_Eventsfeed (GraphQL/SSE, ~30s cadence). Different document shape (CVR_VirksomhedvsVrvirksomhed), person data (CVR_CVRPerson) gated behind a separate OAuth + MitID Erhverv grant we do not hold. The legacy Datafordeler CVR channels deprecate 2027-01-15; the entity-based events model (introduced late 2025) is the go-forward surface there.
The prephase DATAFORDELER/README.md named Datafordeleren the primary bulk
source; the later schema mapping (lassox docs 01–03, 2026-06-26) and the
project charter reversed that in practice. This ADR records the reversal as a
decision.
Three further questions had to be settled for the first implementation: how deltas flow after the initial backfill, what shape the data takes in the hub, and what population is ingested.
Decision¶
distribution.virk.dkis the sole v1 surface for the CVR ingester — backfill and daily delta. One credential, one document shape, one mapper. DatafordelerCVR_Eventsis the documented upgrade path if sub-day freshness is later justified (ownership-change monitoring is the likely trigger), gated on evidence per the posture of prephase ADR-0006.- Refresh is incremental with periodic full reconciliation.
- One-time full-scroll backfill per document type, loaded via a staging table + atomic swap.
- Daily delta job re-scrolls documents with
sidstOpdateret≥ the watermark of the last successful run and upserts transactionally. - A monthly full reconciliation scroll repairs drift (missed deltas, upstream deletions). This is the DDIA rebuild path in scheduled form.
- Every run is idempotent: re-running converges to the same rows, and each
run stamps a run-ledger row (
ingestion_runs) carrying the watermark. - Near-source storage, raw only where cheap.
companiesandproduction_unitsrows carry the full upstream document in arawJSONB column alongside promoted, indexed filter columns.people/people_rolesstore minimized columns only — no raw document (see ADR-0006). Each run's NDJSON extract additionally lands gzipped in GCS as run-level provenance (the deltager extract is minimized before it is written). - Population is all-time, all rows (~2.26M companies including ceased
ones), not the 500k-active slice the prephase sizing assumed. Active is a
filter (
ended_at IS NULL), not an ingestion gate: role and ownership edges reference dead companies, and company-history products need them.
Options considered¶
Delta mechanism¶
| Option | For | Against | Verdict |
|---|---|---|---|
ES delta re-scroll on sidstOpdateret (chosen) |
Same surface, credential, shape, and mapper as backfill; meets the 1-day freshness contract | Daily cadence only; polling | Chosen |
| Full daily re-scroll | Literal prephase ADR-0005; deletions free | ~6.9M docs/day against an ES 1.7 endpoint of unknown tolerance; hours of runtime | Rejected |
CVR_Events as trigger + ES re-fetch |
Near-real-time capable | Two adapters, two credentials, SSE subscription state from day one — no product need yet | Deferred (upgrade path) |
CVR_Events + GraphQL data fetch |
Datafordeler-native | Second complete field mapping; person entities gated behind an OAuth grant we lack | Rejected |
Storage shape¶
| Option | For | Against | Verdict |
|---|---|---|---|
raw JSONB on typed rows, minimized people (chosen) |
Downstream consumers get near-source shape from the hub; one row per entity; filters stay on indexed columns; PII surface bounded | Tens of GB in Postgres → Supabase compute add-on (priced in prephase ADR-0002) | Chosen |
| Promoted columns only, raw in GCS | Smallest hub | Starves downstream scenarios of source shape; every new need becomes a re-mapping round-trip | Rejected |
| Separate replication tables + typed projections | Cleanest derived-data separation | Two copies of everything and a second load stage for a solo operator | Rejected |
Consequences¶
- Cleartext credential constraint: because the surface is HTTP-only with Basic Auth, ingestion runs only from the Cloud Run batch runtime (GCP egress), never from developer machines on untrusted networks. The credential is treated as low-trust and rotatable.
- ES 1.7 dialect is a contract. Modern client libraries and query syntax cannot be assumed; the adapter pins its own query bodies and is tested against recorded fixtures of real responses.
- Upstream dependency concentration. One surface means one failure
domain; a
distribution.virk.dkoutage stalls freshness until it returns. Accepted under the 1-day contract; the reconciliation job doubles as catch-up. - Watermark discipline. The delta query only works if the watermark advances solely on successful runs; a failed run must not move it.
- Deletion visibility is monthly, not daily: upstream deletions surface
only at reconciliation. Acceptable because CVR expresses most lifecycle
changes as status updates (
SLETTET,UNDER KONKURS), not row removal. - Storage sizing moves from "500k skinny rows" to "millions of rows with raw documents": the Supabase compute add-on decision in prephase ADR-0002 becomes live at backfill time, not at 20-source scale.
Amendment 2026-07-10 — CVR_Events trigger contract¶
The "trigger metric" open question below is resolved into a recorded,
measurable adoption condition (research:
gap analysis G9/G10).
Adopt CVR_Events — in the shape this ADR already chose, events as trigger
plus ES re-fetch — when either holds:
- Product clause: a committed product scenario requires detecting a CVR change class in under 24 hours (faster than the daily delta can deliver); ownership-change alerting remains the expected case.
- Evidence clause: the daily delta measurably breaches the 1-day freshness contract — systematic misses found by two consecutive monthly reconciliations, or observed delta lag exceeding the contract, as reported by the ops-health work (E4).
Until a clause fires, no Datafordeler events adapter is built. The first task
when one fires is a docs/data-sources/ access test of the entity-based
events surface (schema, cursor/replay, delivery guarantees, person-event
visibility without the OAuth grant). The 2027-01-15 legacy-channel
deprecation does not affect this repository's chosen surface and is not a
forcing function.
Amendment 2026-07-19 — staging + atomic swap loader implemented (#20)¶
The backfill loader described in the Decision section is implemented:
bulk_load_companies (ingesters/cvr/loader.py) creates a companies_staging
table (LIKE companies INCLUDING ALL), COPYs the full scroll result into it,
then swaps it in as companies via two ALTER TABLE ... RENAME statements —
all inside one transaction that only commits once both renames succeed. A
reader against companies therefore only ever sees the complete old
population or the complete new one; a failure at any point (including the
upstream scroll itself raising) leaves companies untouched. It refuses to
swap in an empty result rather than silently truncating the table, since a
real full scroll returning zero documents means the scroll failed, not that
the population is empty. Proven with a Postgres-backed test suite
(tests/test_cvr_loader.py, tests/test_cvr_pipeline.py), not yet against
the live upstream — see the throughput open question below, which this
amendment does not close.
The backfill entrypoint (ingesters/cvr/__main__.py) always uses this
staging+swap loader. It once selected between two loaders on --max-docs,
using a transactional upsert for size-capped tracer-bullet/dev runs (#19);
that flag and its loader were deleted on 2026-08-16, since a run is restricted
by bounding the population (#248) and a bounded run still swaps. Since #720,
the swap preserves rows outside that range (see the amendment below). The scheduled
Cloud Run Job
(infra/batch_job.tf) is deliberately left capped for now — an uncapped run
is a one-time operator action (gcloud run jobs execute ... --args=...
overriding the default args), not a daily schedule; the daily slot moves to
the delta job once #21 ships, per this ADR's incremental-refresh decision.
Amendment 2026-09-10 — bounded replacement preserves unexamined records (#720)¶
The owner approved replacing only the selected CVR range after reconciliation run 71 tried to remove CVR 41527080 and failed on its financial-report foreign key. A range is the scope of a rebuild, not authority to remove other data. Companies, production units, and financial reports keep complete outside rows, including their original cache timestamps. NULL-CVR rows are also preserved unless the source explicitly returns a new version. An observed production unit that moves into the range replaces the outside row with the same primary key; a conflict on any other unique key aborts the swap. The staged population includes these preserved rows before dependent indexes, auditor publications, and financial-metric cleanup run. Loaded-row counts report only newly ingested rows. Empty-input and catastrophe guards still apply; row and null-rate checks compare only the selected range, so outside rows cannot hide a broken scroll. Unbounded runs still replace the full population.
Participant identities are global, not owned by one company. A bounded participant scroll updates the returned identities and replaces role assertions only inside its range. It does not delete absent global identities: those people may still serve unexamined companies. The deletion ceiling is recorded as skipped with this reason. Full-register reconciliation retains the global deletion pass.
Amendment 2026-07-19 — daily delta job implemented (#21)¶
The delta half of the incremental-refresh decision is implemented: run_delta
(ingesters/cvr/pipeline.py) reads the watermark of the last successful run
(ledger.get_last_watermark), re-scrolls only documents with sidstOpdateret
at or after it (sidst_opdateret_since_body), and upserts them via the
existing transactional loader — never staging+swap, since a delta result is
a subset of the population, not the whole thing. The watermark only advances
on success (existing ledger discipline); a failed delta leaves it untouched
and is safely re-runnable. With no prior successful run it falls back to the
Unix epoch, so a fresh companies table can be brought up via delta alone
without a hard ordering dependency on backfill having run first — at the
cost of that first delta scrolling the entire population once, same as a
backfill would, just through the upsert path instead of staging+swap.
infra/batch_job.tf's daily schedule now runs delta by default, per the
"daily slot moves to the delta job" note above; the one-time full backfill
is the manual gcloud run jobs execute ... --args=... operator action
already documented there. Proven with a Postgres-backed pipeline-seam test
suite (tests/test_cvr_pipeline.py) covering: changed rows upsert and the
watermark advances; the query sent upstream is bounded by the last
watermark; the epoch fallback when no watermark exists; idempotent
convergence on repeat; and a failed run neither advances the watermark nor
touches existing rows. Not yet run against the live upstream — same
Cloud-Run-only constraint as the backfill amendment above, so it doesn't
close the throughput open question either.
Amendment 2026-07-19 — production units on the shared framework, FK-safe swap (#22)¶
The second entity — production_units, 1:N to companies via cvr — landed
on the framework built for companies with no ingestion code forked:
ingesters/cvr/pipeline.py now takes an EntityConfig (ES type, table, PK
column, columns, mapper) and _run/_run_delta are entity-agnostic;
ingesters/cvr/loader.py's upsert_rows/bulk_load_rows are the same
generalization for the loaders. Only the migration, the VrproduktionsEnhed
mapper, and CLI --entity wiring are new, per the issue's acceptance
criterion. production_units reuses raw_store, ledger, and the
cvr_es adapter completely unchanged.
This surfaced a real correctness gap in the #20 staging+swap loader,
found by the loader's own test suite once the production_units.cvr FK
existed: a rename-based swap doesn't carry other tables' foreign keys to the
newly named relation — Postgres binds an FK constraint by OID, not by name —
so swapping companies left production_units.cvr still pointed at the
old, about-to-be-dropped table and blocked the drop outright
(DependentObjectsStillExist). The reverse direction is also broken:
CREATE TABLE ... LIKE ... INCLUDING ALL never copies a table's own
outgoing FK constraints regardless of INCLUDING options, so swapping
production_units itself silently produced a new table with no FK to
companies at all. bulk_load_rows now captures every FK constraint
touching the table being swapped (either direction) before starting, and
re-adds each one verbatim (via pg_get_constraintdef) once the swap
completes — which also means a full backfill that would genuinely orphan a
child row now fails loudly with ForeignKeyViolation instead of silently
losing referential integrity. This generalizes for any future FK the hub
adds (e.g. people_roles -> companies/people), not just this one.
Per the plan's open question: a P-unit's cvrNummer parent can change
(unit reassigned to a different company); upserts keyed on p_nummer
overwrite the FK rather than orphaning or duplicating — proven directly by
a pipeline-seam test with the same p_nummer reassigned to a new cvr.
infra/batch_job.tf gets a second Cloud Run Job + Cloud Scheduler entry
(cvr-ingester-production-units), same image/identity/secrets as the
companies job, running delta --entity production-units daily — one
Cloud Run Job per scheduled entity is the normal infra shape here, distinct
from the "no framework forked" rule which is about the ingestion code path.
Not yet run against the live upstream — same Cloud-Run-only constraint as
the amendments above.
Amendment 2026-07-19 — people & roles, PII-minimized (#23)¶
The M:N core of the hub — people + people_roles from the deltager
type — is implemented per ADR-0006's minimization contract. This entity
doesn't fit the flat-row EntityConfig shape the last two amendments
reused: one deltager document maps to a person row plus N role rows
(one per company relation × organisation), so it gets its own pipeline
functions (run_people_backfill/run_people_delta /
ingesters/cvr/pipeline.py's _run_people) rather than reusing
_run/_run_delta — while still reusing core/ledger.py,
core/raw_store.py, core/db.py, and the cvr_es adapter completely
unchanged, same as the other entities.
Minimization is enforced by construction, not by filtering:
map_deltager only ever reads the allowlisted fields it promotes to
columns (registry surrogate key, current name, participant kind, and
role/ownership data) — residential addresses and every other non-role
attribute are never referenced, so there's nothing to accidentally leak.
minimize_deltager_for_extract rebuilds the document from that same
allowlist for the GCS extract, so the run's raw-provenance copy can never
carry more than the row mapping does. Mapper tests assert the absence of
the dropped fields (not just presence of kept ones), per ADR-0006's
explicit testing requirement, and a pipeline-seam test asserts the same
against the actual (gzipped, decompressed) extract object.
Role replacement is delete-then-insert per participant
(loader.replace_person_roles), not upsert or staging+swap: each
deltager document is authoritative for that participant's entire
current role set, so a role absent from the latest document (because it
quietly ended upstream) must be dropped, not left stale — proven by a
pipeline-seam test. This is also why people/roles never needs a
staging+swap loader variant at all: unlike a full-population table replace,
a capped or delta run of deltager only ever touches the participants it
actually scrolled, which is already safe by construction. One consequence
worth naming: each participant's replace commits independently (not the
whole batch in one transaction), so a mid-batch failure leaves
already-processed participants durably persisted while the ledger records
the run as failed and the watermark doesn't advance — safe because
replace_person_roles is itself idempotent, so the next run (from the
unchanged watermark) simply reprocesses the same range.
Roles referencing ceased companies load correctly: a ceased company is
still a row in companies (this ADR's all-time-population decision), so
people_roles.cvr's FK holds without special-casing.
infra/batch_job.tf gets a third Cloud Run Job + Scheduler
(cvr-ingester-people), running delta --entity people daily. Unlike
companies/production-units, there's no separate one-time full-backfill
operator step: delta's epoch fallback already self-bootstraps a fresh
people/people_roles population on its first run, since deltager never
needed the staging+swap path in the first place. Not yet run against the
live upstream — same Cloud-Run-only constraint as the amendments above.
Amendment 2026-07-19 — monthly reconciliation, diff-vs-rebuild decided (#24)¶
Diff-vs-rebuild decision (closing the open question below), decided by architecture rather than observed drift volume — no live run has happened yet to observe real drift, so this is a principled choice, not a volume-driven one; observed numbers remain the open item they always were:
- Companies / production units: full staging+swap rebuild.
run_reconcileis literallyrun_backfillwithrun_type="reconcile"instead of"backfill"— the swap already replaces the entire population, which repairs drift (every row gets fresh data) and deletions (a row absent from the new scroll isn't in the swapped-in table) for free. Since the staging+swap infrastructure already exists and is already proven (#20), reusing it for reconciliation is strictly cheaper than building a second, parallel diff-and-repair path for these two entities — there is no drift-volume threshold that would change this: rebuild is never more expensive than diff-and-repair here, since the rebuild is the existing backfill. - People/roles: diff-and-repair.
deltagernever uses staging+swap (replace_person_rolesis delete-then-insert per participant, for both backfill and delta) — there's no swap infrastructure to reuse for reconciliation, so drift-repair falls out of the same per-participant replace a reconcile's full scroll already does, and deletion-repair needs an explicit pass:loader.delete_stale_peoplestages everyperson_idseen in the reconcile scroll into a run-scoped staging table and deletes anypeople/people_rolesrow not in that set (anti-join, not aNOT INlist, since a real reconcile scrolls ~1.82M participants). Refuses to run against an empty seen-set, mirroringbulk_load_rows's empty-population guard.
Proven with a Postgres-backed pipeline-seam test suite seeding both a
drift case (a row diverged from what the fixture scroll would produce) and
a deletion case (a row absent from the fixture scroll) for companies,
production_units, and people — reconcile repairs both in every case,
tagged run_type = 'reconcile' in the ledger, and is idempotent
(re-running twice converges to the same counts).
Concurrency safety ("reconciliation is safe to run concurrently-adjacent to a daily delta" acceptance criterion) is a property of mechanisms already in place, not new code: companies/production_units reconcile's swap runs in one uncommitted transaction, so per Postgres MVCC a concurrent delta's session only ever sees the complete pre-swap table or the complete post-swap one, never an in-between rename state (the same guarantee #20 already established for "never observes a partial population") — a concurrent statement that needs a lock the swap holds simply waits for it, which is blocking, not a deadlock, since neither side waits on a lock the other is waiting for in return. For people, reconcile and delta both act per-participant (delete-then-insert / replace), so they only conflict if they touch the exact same participant at the same instant — an acceptable, narrow last-writer-wins race given both operations are idempotent and re-runnable. The 2026-09-01 live run disproved the earlier reasoning for multiple swaps: foreign-key DDL created a cross-table deadlock even though each individual swap was transactional. The amendment below records the advisory-lock correction and its lock-contention regression test.
infra/reconcile_jobs.tf adds one Cloud Run Job + Scheduler per entity
(cvr-ingester-reconcile-{companies,production-units,people}), same
image/identity/secrets as the daily jobs, on a monthly schedule
(var.reconcile_schedule, default 0 4 1 * *). Kept in a separate
for_each-driven file rather than folding the existing three daily
job/scheduler pairs into a loop, since CD applies infra/ on every merge
to main — refactoring already-live resource addresses would force
Terraform to destroy and recreate them for no functional gain. Not yet run
against the live upstream — same Cloud-Run-only constraint as every
amendment above.
Amendment 2026-09-01 — reconciliation pause and swap serialization (#408)¶
The first live monthly run showed that separate transactional swaps can still deadlock: each swap drops and recreates foreign keys, and the four schedules started in the same minute. A transaction-scoped PostgreSQL advisory lock now serializes the swap phase across all flat-row entities. It does not rely on a session lock, because the hosted database uses transaction pooling.
Monthly reconcile schedulers have their own reconcile_schedulers_paused
control, defaulting to true. A bounded daily delta can remain enabled while
monthly full-population repair stays paused until the operator has a coherent
population and approves the ordered run. The existing
ingestion_schedulers_paused control still pauses all ingestion schedulers
when it is set to true.
People reconciliation also no longer uses a temporary seen-person table. Its run-scoped durable staging table survives the commit after each scroll batch, which is required by transaction pooling. Staging rows are deleted after the run and stale rows are cleaned at the next reconcile start.
Amendment 2026-07-19 — read contract shipped, M1 filter measured (#26)¶
The consumer-facing surface is live: active_companies (a filtered view
over companies) plus a composite partial index
(companies_m1_filter_idx on (main_industry, municipality_code,
employees_band) where ended_at is null) for the M1-style filter — see
docs/reference/read-contract.md for the full contract, versioning
posture, and PostgREST exposure model. Every base table gets row-level
security enabled with no policies, which blocks any non-owner role
regardless of table-level grants; only the view is granted to
authenticated.
Unlike the throughput/drift-volume open questions above, this one is
measured for real, not deferred: generating a realistic-scale population
needs no live upstream call, just synthetic data, so
tests/test_read_contract_performance.py (opt-in, RUN_PERFORMANCE_TESTS=1)
loads 2,260,000 synthetic rows — companies' own population estimate from
this ADR's Decision section — and measures the M1 query for real: 0.25ms
execution time via an Index Scan on the composite index, comfortably
under the <1s target.
This surfaced a third instance of the swap-dependency problem first
found in #22 (foreign keys) — views bind to their base table by OID too,
so a naive swap would leave a dependent view reading the dropped old table
and block the drop the same way an unhandled FK did. Row-level security
enablement isn't carried by LIKE ... INCLUDING ALL either (neither are
views, at all — Postgres's LIKE clause has no INCLUDING VIEWS option to
begin with). bulk_load_rows now captures dependent views (definition +
grants) and the table's RLS state before every swap, and re-establishes
both afterward — a swap without this would have silently reopened
companies to any role with a stray grant (RLS disabled) and broken
active_companies entirely (dangling view). Generalizes for any future
view or RLS policy on a swapped table, not just this one.
Cosmetic, not functional, side effect confirmed by the same measurement
run: the composite index's migration-given name
(companies_m1_filter_idx) doesn't survive a swap — Postgres's
LIKE ... INCLUDING INDEXES regenerates index names from their columns,
not the original name (already tolerated by #20's own
index-preservation test, which asserts a substring rather than an exact
name for exactly this reason). The index's columns, partial predicate, and
the query plan choosing it are unaffected — confirmed by the EXPLAIN
output in the performance test above still showing an Index Scan, just
under the auto-generated name. Not fixed: correlating pre/post-swap index
identity by definition rather than name would be a real generalization
(same shape as the FK/view fixes) but for a cosmetic-only concern; revisit
if a future need (e.g. monitoring pg_stat_user_indexes by name) makes the
name itself load-bearing.
Amendment 2026-07-19 — operational health: staleness view + alerting live (#25)¶
Closes out the CVR ingester epic (#15) — every remaining E-series slice is now delivered.
Staleness view: ingestion_staleness always returns exactly one row
per known entity (virksomhed/produktionsenhed/deltager), even ones
that have never run — a left join against a fixed three-row values
list rather than a plain group by, since a fresh database with an empty
ingestion_runs would otherwise silently report nothing instead of
flagging every entity as unmonitored. stale = no successful run within
the 1-day freshness contract (this ADR's Decision section) plus a 6h
margin against normal daily-cadence variance, or no successful run ever.
Deliberately keyed on last successful run time, not watermark age: a
delta run that succeeds but finds zero upstream changes correctly doesn't
advance the watermark (existing discipline), so watermark age alone would
false-positive on a quiet day — execution recency is what "is the pipeline
still running" actually asks.
Alerting reuses one mechanism for both failure classes. Rather than
building a second, staleness-specific alerting path (e.g. a log-based
metric), ingesters/cvr/__main__.py check-staleness queries the view and
exits non-zero on any breach — which means a stale pipeline and a failed
one both produce the same Cloud Monitoring signal
(completed_execution_count{result="failed"}), so one alert-policy shape,
applied to every scheduled job including the staleness check itself,
covers both of this issue's notification acceptance criteria. Wired via
infra/monitoring.tf (activated from the bootstrap
monitoring.tf.example template, not hand-configured — no
deploy_monitoring opt-in flag; it's on by default once this merges) →
infra/staleness_check_job.tf (a new hourly Cloud Run Job) →
.github/workflows/gcp-incidents.yml (schedule trigger added, pointed at
the Pub/Sub subscription's deterministic name since a scheduled trigger
has no inputs to read a value from) → GitHub issue.
Proven with pipeline-seam tests directly manipulating ingestion_runs'
finished_at to simulate fresh/old/never-run states (Postgres, not
mocked) and an end-to-end CLI smoke test against the same local instance
confirming both the zero-breach success path (exit 0) and the breach path
(exit 1, breached entities printed). Not proven: an actual Cloud
Monitoring alert reaching a real GitHub issue, which needs a live Cloud
Run execution — same operator-follow-up limitation as every other
amendment in this ADR.
Open questions¶
- Real scroll throughput and rate tolerance of
distribution.virk.dk, still open: the loader that would carry a full ~6.9M-doc backfill exists (#20 amendment above), but no live run against the real upstream has happened yet — the dev sandbox this was built in cannot reachdistribution.virk.dk(Cloud-Run-only per this ADR's cleartext-Basic-Auth consequence). First Cloud Run execution of the uncappedbackfillcommand in test should record observed throughput/rate-limit behavior here and close this question for real. - ES version/dialect verification: the unauthenticated cluster root
reports Elasticsearch 6.8.23 (live probe 2026-07-10), contradicting the
ES 1.7 guidance this ADR's "dialect is a contract" consequence rests on.
Verify modern-DSL support (
search_after, sliced scroll) against the authedcvr-permanentindex — the same first live Cloud Run execution above is the first chance to observe this — then correctdocs/data-sources/cvr-elasticsearch.mdand this consequence. - ~~Whether the monthly reconciliation is a diff-and-repair or a full staging+swap rebuild~~ — resolved by architecture in the 2026-07-19 reconciliation amendment above (rebuild for companies/production_units, diff-and-repair for people/roles); real observed drift volume is still pending the same live-run constraint as the throughput question above.
- ~~Trigger metric for adopting
CVR_Events~~ — resolved by the 2026-07-10 amendment above.
Amendment 2026-08-12 — production-unit parent and people-role identity (#175, #178)¶
VrproduktionsEnhed.virksomhedsrelation is a historised list, not a flat
object. The mapper selects the entry whose periode.gyldigTil is null, then
falls back to the final entry for a ceased unit, which is the established CVR
history-selection rule. A unit with an empty relation list remains in
production_units with a null cvr: the source still identifies a real
production unit, and the column has always been an optional FK. This preserves
the raw source record without fabricating a parent company.
people_roles.valid_from is source knowledge, not an identity component:
CVR can omit it. people_roles therefore has a surrogate primary key and a
unique nulls not distinct (cvr, person_id, role_type, valid_from) source
identity. The unique constraint preserves one undated role edge across replay;
the surrogate key avoids claiming that an absent date can participate in a
primary key. The existing per-participant replace remains the idempotency
boundary.
Amendment 2026-08-12 — raw JSONB leaves Postgres; GCS is the only raw store¶
This amendment reverses decision 3's "raw JSONB on typed rows" for
companies and production_units, and with it the "Promoted columns only,
raw in GCS" row of the Storage-shape options table, which this ADR rejected
on the argument that it "starves downstream scenarios of source shape". The
argument was made against an unmeasured cost. The cost is now measured and it
is the dominant term in the whole hub.
Decision. Drop the raw column from companies and production_units.
Promoted columns stay exactly as they are. The run-scoped gzipped NDJSON
extract in GCS — already written on every backfill, delta, and reconcile run
by raw_store.write_ndjson_gz — becomes the single raw-snapshot of record.
No compatibility path, no dual-write period: the column is dropped in one
migration (project rule 2).
Evidence. Measured 2026-08-12 by loading live sampled documents through the real mappers into this repository's real schema, then scaling by live population counts (not estimated — an earlier estimate put the hub at 28 GB and was wrong in both directions):
| table | with raw |
without raw |
raw share |
|---|---|---|---|
companies (2,262,234 rows) |
3,568 B/row → 8.07 GB | 217 B/row → 0.49 GB | 94% |
production_units (2,878,879 rows) |
1,836 B/row → 5.29 GB | 213 B/row → 0.62 GB | 88% |
A company row is 217 bytes of business data wrapped in ~3,050 bytes of
pglz-compressed source document. The mean live Vrvirksomhed document is
8,486 bytes of JSON; Postgres TOAST compresses it to ~3,050 bytes, while gzip
in GCS reaches 613 bytes for the same content. GCS stores the identical
document about 5x more efficiently than Postgres can, which makes holding it
in both places not merely redundant but redundant at a premium.
The "starves downstream scenarios" argument also failed its own test: no view,
no read contract, and no Python query reads the raw column. It is
write-only. Every raw reference in src/ is either a loader column list, a
mapper input, or the GCS raw-store — verified by grep at the time of writing.
Consequences.
- The prephase ADR-0002 Supabase compute add-on, which the original Consequences section declared "becomes live at backfill time", is deferred again. The four-entity hub lands near 2.6 GB rather than 22 GB.
- Mapping-bug recovery changes shape, and this is the real cost. Today a
bad mapping is repaired by re-reading
rawin SQL. After this change it is repaired by re-reading batch NDJSON objects from GCS. Those objects are immutable and complete, but they are keyed by run and batch, not indexed by CVR or p-nummer, so single-entity recovery is a scan rather than a lookup. Accepted: bulk re-mapping reads whole batches anyway, and the live ES surface answers any single-entity question directly. - The GCS extract stops being belt-and-braces provenance and becomes load bearing. Bucket lifecycle rules, retention, and deletion protection now guard the only copy. Any future change to the raw-store write path is a change to the system of record.
- Reversal is cheap: re-adding the column and repopulating it from GCS is a batch job, not a re-scroll of upstream.
Amendment 2026-08-12 — source documents live in GCS, not in Postgres (#134)¶
The original decision kept the complete upstream document in a raw jsonb
column beside every row, on the reasoning that a mapping mistake should be
recoverable without re-scraping. The recovery property was right. The storage
location was wrong.
Measured by loading 500 real rows per table: raw is 3,048 bytes of a
3,949-byte companies row — 77% of the table, against 171 bytes for every
promoted column combined. Across the populations that is 27.3 GB, of which
16.7 GB is raw. Nothing read it: no view, no read-contract query, no
application code. The only references in the repository were the three column
definitions and the mappers that wrote it.
Meanwhile the identical documents were already being written to GCS as immutable gzipped NDJSON snapshots, before every load, at roughly 682 compressed bytes per company. The column was a second copy of the system of record, kept in the most expensive place available.
Source documents therefore live in the raw store only. Each row carries
raw_snapshot_uri, naming the GCS object its current values were mapped from,
at roughly 76 bytes instead of 3,048 — so "which snapshot produced this row?"
stays answerable per row, which it was not before. A delta upsert moves the
pointer forward, because it names the object that produced the values in the
row now, not the first object that ever created it.
What is given up, explicitly: reading a source document is a GCS fetch of one batch object rather than a SQL column, and those objects are keyed by run and batch rather than by CVR. For a single company the live CVR API answers faster than either. The snapshots keep doing what they were always for — recording what the source said at a point in time — and the mapping-recovery property is unchanged, since a re-parse reads the snapshot rather than re-scrolling upstream.
deltager is unaffected: it never had a raw column, because ADR-0006
minimises personal data before anything is persisted anywhere, including the
snapshot.
Population checkpoint and freshness amendment (#782)¶
A checkpoint covers exactly the population label recorded by its run. A new
population starts from EPOCH; it does not use another population's checkpoint.
This deliberately avoids range-containment inference. Runs without a watermark
never supply a checkpoint. The operator view ingestion_staleness describes
the full population only. ingestion_freshness(population_bound) checks one
explicit population, and the scheduled check receives the same CVR range as
the scheduled ingesters. A bounded deployment can therefore be fresh without
claiming that the full population is fresh.
Population publication amendment (#782)¶
Prepare each population in a uniquely named table. Release the live table's shape-copy lock before reading the remote source. Build indexes after COPY. At publication, acquire the shared swap lock and the live table lock, then check the original relation, column shape, and statement-triggered write revision. A concurrent write or schema change rejects publication; the operator retries acquisition. This policy does not discard a concurrent delta.
Foreign keys, views, triggers, and derived projections remain in the atomic
publication transaction. DDL and validation still take locks. The loader records
staging, lock wait, and publication durations separately. Normal failures remove
staging tables. After a killed process, an operator can remove its uniquely named
*_stage_* table after confirming that the owning job has stopped.