Ingestion quality gates: typed record contracts, record quarantine, load guard rails, and a per-run quality verdict¶
Status: accepted Date: 2026-08-10 Deciders: Chris (solo founder) Depends on: ADR-0005 (run ledger, watermark discipline, staging+swap, raw-store rule), ADR-0006 (minimization before persistence), ADR-0007 (document fetch, versioned extractor, "missing ≠ zero"), ADR-0009 (execution durability is a separate concern) Scope: registry ingesters (CVR, Regnskabsdata). Enrichment pipelines inherit the framework; their own gates are a later decision.
Context¶
ADR-0005 gave this repository a genuinely strong durability story, and
ADR-0007 reused it unchanged: an ingestion_runs ledger, watermarks that
advance only on success, immutable raw snapshots in GCS, and staging+atomic
swap so a reader never sees a partial population. ADR-0009 is separately
spiking DBOS for execution durability — resuming an interrupted process.
None of that is data quality. Durability answers "did the run finish, and can it be replayed?". It does not answer "was what the run loaded correct?". Reviewing the shipped ingesters (CVR #16–#26, Regnskabsdata #54–#58) against that second question surfaced five gaps, none of which DBOS or the existing ledger addresses:
G1 — No population-level guard on the swap. bulk_load_rows refuses to
swap in a zero-row result, on the reasoning that a real full scroll never
returns zero documents. But the same reasoning applies with equal force to a
scroll that returns 40% of the population — a truncated scroll, an upstream
partial outage, a silently-expired scroll cursor. Today that result swaps in
atomically, deletes 60% of companies, stamps status = 'succeeded', and
advances the watermark. delete_stale_people has the same shape. This is the
highest-blast-radius failure the system currently permits, and it is
undetectable after the fact except by noticing the row count.
G2 — No structural contract at the ingest boundary. pydantic is a
declared dependency and is used nowhere in src/. The mappers
(ingesters/cvr/mappers.py, ingesters/regnskab/mappers.py) read raw dicts
with two mixed failure modes:
- Required-field access raises
KeyError—doc["Vrvirksomhed"],metadata["nyesteNavn"]["navn"],regnskabsperiode["startDato"],doc["sidstOpdateret"]. One malformed document out of ~1.3M companies or ~6.4M publications aborts the entire run. - Everything else is
.get()with aNonefallback. If an upstream field is renamed or restructured, the whole population silently loads as NULL, the run reports success, and the watermark advances. There is no artifact anywhere in the repository that declares which upstream fields an ingester actually requires.
The second mode is the more dangerous: it is silent, it passes every existing gate, and the ledger records it as a healthy run.
G3 — Failure granularity is the whole run. Anything that raises inside a document loop rolls back and fails the run. There is no quarantine, so a run cannot report "1,299,998 loaded, 2 rejected". A single poison record wedges a delta indefinitely: it fails, holds the watermark, is retried on the next schedule, fails identically, and only surfaces when the staleness check breaches ~30 hours later. Nothing points the operator at the offending document — it is somewhere in the run's raw NDJSON, unindexed.
G4 — No semantic validation, and absence is untyped. Nothing asserts
period_start <= period_end, that a CVR number is eight digits, that
ownership_pct is in 0–1, or any financial-statement coherence. Related:
xbrl_extractor._extract_concept returns candidates[0] if len(candidates) == 1
else None — an ambiguous fact becomes NULL with no record that ambiguity
occurred. ADR-0007 correctly established "missing ≠ zero", but the only
coverage signal is financial_metrics row presence, which conflates not
filed digitally, PDF-only, outside the backfill window, fact absent from
the filing, ambiguous context, and parse failed. A consumer cannot
distinguish "this company files no turnover" from "our extractor could not
read it".
G5 — Retries do not distinguish retryable from terminal.
CvrEsClient._post and DocFetcher._get_with_retry both catch
requests.RequestException, which raise_for_status() raises for 4xx as well
as 5xx. A 401 or a 404 is retried to exhaustion (five and three attempts
respectively) at linear backoff with no jitter, and Retry-After is ignored.
Wasted attempts against an upstream we are explicitly obliged to treat
politely (ADR-0005, ADR-0007).
Underneath all five is the operator-facing problem: the ledger records
counts, not quality. Answering "was last night's run good?" means querying
ingestion_runs and eyeballing rows_upserted against intuition. There is no
per-run record of what was checked, what the measured value was, what the
threshold was, and whether it passed.
Decision¶
1. Every source document passes a typed record contract before mapping¶
Each source entity gains a record contract: a pydantic model declaring the shape an ingester requires of one upstream document. Contracts live beside their mapper and are the mapper's only input — mappers stop reading raw dicts.
Strictness is deliberately asymmetric:
- Required and strictly typed: identity fields, the watermark field, and
every field backing a target column declared
not null. ForVrvirksomhedthat iscvrNummer,enhedsNummer,sidstOpdateret, andvirksomhedMetadata.nyesteNavn.navn; foroffentliggoerelse, the ES_id,sidstOpdateret, and bothregnskabsperiodedates. - Optional and nullable: every other promoted field. These do not fail a record. They are instead watched at population level by null-rate guard rails (decision 4), which is what actually catches an upstream rename.
extra="allow". Unknown upstream fields are tolerated, not rejected. We do not control the registry's release cadence, andextra="forbid"would convert every cosmetic upstream addition into a full-population outage.
Contracts are versioned (CONTRACT_VERSION per entity, stamped on the
run) for the same reason the XBRL extractor is: a contract change alters what
we accept, and that must be attributable after the fact.
This is structural validation only. It answers "is this document shaped the way we require?", never "is this value plausible?".
2. A record that fails a gate is quarantined, not fatal¶
New ingestion_rejects table, one row per rejected source document, holding:
run_id, source, entity, the source key where one is recoverable, the
stage that rejected it (fetch, contract, map, semantic, load,
parse), a machine-readable reason_code, the human-readable detail, and the
raw document as jsonb.
The raw document is stored subject to ADR-0006: deltager rejects persist the
minimized projection, never the full participant record. Minimization
happens before quarantine, exactly as it already happens before the GCS
extract.
A rejected record does not fail its run. The run continues, loads every valid record, and reports both counts.
3. A reject rate breach fails the run and holds the watermark¶
Per-entity reject-rate thresholds are declared config, evaluated at the end of
the record loop and before the load commits. Under threshold, the run
succeeds and the watermark advances. Over threshold, the run fails through the
existing path — rollback, finish_run_failure, watermark held — with the
breach recorded as the error.
This is the decision that makes quarantine safe. Isolated upstream junk stops wedging the pipeline; systemic drift still stops the pipeline, which is the only behaviour that keeps a bad population out of the read contract.
Thresholds start deliberately tight (order of 0.1–1% depending on entity) and are widened only against observed production evidence, recorded in this ADR.
4. Population-level guard rails run inside the load transaction¶
Before any staging+swap commit or bulk delete, the load evaluates guard rails against the prior successful run for the same source/entity:
- Row count. A staged population that has lost essentially all of the live one aborts the swap. The existing zero-row guard is its degenerate case.
- Per-column null rate. A promoted column that is essentially all NULL in the staged population, and was populated in the live one, aborts the swap. This is the gate that catches G2's silent rename, which no per-record contract can catch by construction.
- Deletion ceiling.
delete_stale_peopleand_delete_orphaned_metricsrefuse a run that would remove essentially the whole population.
Amended 2026-08-16 (#251). These were relative bands — a 5% shrink, a 10% jump in a column's null rate. The first time they ever evaluated against real data they refused a legitimate swap, because a bounded CVR block and an unbounded sample are genuinely different populations: 91.3% of the oldest block has no email address against 38.1% of the whole register. Nothing was renamed.
A band asks whether a population is different, and the honest answer is
often yes. These rails now ask whether it is broken, at one absolute
threshold — CATASTROPHIC_SHARE, 99.5% — with no per-entity knobs, because a
catastrophe is a catastrophe in every table. This also disposes of the
"compare like with like" design #251 proposed: a rail that fires only on
catastrophe does not care which population it is looking at.
Guard rails run inside the same transaction as the swap, so a breach leaves
the previous population exactly in place — the property bulk_load_rows
already guarantees for any other failure.
The first run for an entity has no baseline. It passes the delta rails and
records them as skipped (no baseline) rather than silently passing.
5. Semantic rules are a declared, versioned rule set, separate from contracts¶
Semantic validation is a distinct stage from structural validation, with a distinct reason-code namespace, because the two have different failure economics: a contract violation means we cannot read the document, a semantic violation means we can read it and do not believe it.
Rules are declared per entity (period_start <= period_end, CVR is eight
digits, ownership_pct ∈ [0,1] and voting_pct ∈ [0,1], valid_from <=
valid_to, financial sign/coherence rules), each with a severity: reject
quarantines the record, warn loads it and counts the warning into the run's
quality verdict. Out-of-range ownership and voting shares warn and load.
6. Absence gets a typed reason¶
financial_metrics gains a per-metric absence taxonomy so "missing" stops
being one undifferentiated state. The extractor records which of
no_xbrl_document, outside_backfill_window, fact_absent,
ambiguous_context, or parse_failed produced a NULL. _extract_concept's
existing "more than one candidate ⇒ None" branch becomes an explicitly
recorded ambiguous_context, not a silent NULL.
This makes ADR-0007's "missing ≠ zero" rule enforceable by a consumer rather than merely documented, and it turns extractor quality into something measurable per taxonomy version.
7. Retries classify errors before retrying¶
A shared retry policy in core/ replaces the two ad-hoc loops. Retryable:
connection errors, timeouts, 429, and 5xx. Terminal: every other 4xx, raised
immediately. Backoff is exponential with jitter, honours Retry-After when
present, and is bounded by a total time budget as well as an attempt count.
Both CvrEsClient and DocFetcher adopt it; retry attempts and terminal
classifications count into the run's quality verdict.
8. Every run publishes a quality verdict¶
The operator-facing artifact, and the point of the whole ADR. New
ingestion_run_quality table: one row per run per gate, recording the gate
name, its outcome (pass / fail / warn / skipped), the measured value,
the threshold it was compared against, and the contract/rule-set version in
force. A ingestion_run_report view joins it to ingestion_runs so one query
answers, per run: what ran, what it read, what it loaded, what it rejected and
why, which gates were evaluated, and which of them held.
Gates that did not run are recorded as skipped with a reason. A gate's
absence from a run's verdict is itself a finding, not a silent pass.
Options considered¶
Record-level failure semantics¶
| Option | For | Against | Verdict |
|---|---|---|---|
| Quarantine + rate threshold (chosen) | Poison records stop wedging deltas; systemic drift still halts the pipeline; rejects are indexed and queryable | Two mechanisms to reason about; thresholds need calibration against production evidence we do not have yet | Chosen |
| Fail the whole run (status quo) | Simplest; strictly safe; no new table | One malformed document out of millions wedges a delta until a human intervenes ~30h later, with no pointer to the offending record | Rejected |
| Quarantine, never fail | Maximum availability | Systemic upstream drift lands silently as a large reject pile with no forcing function; converts a hard gate into a dashboard nobody is obliged to read | Rejected |
Contract strictness¶
| Option | For | Against | Verdict |
|---|---|---|---|
Required fields typed, rest tolerant, extra="allow" (chosen) |
Catches the failures that matter (missing identity, missing watermark, not null violations) without coupling us to a registry release cadence we do not control; optional-field drift is caught at population level instead |
Does not detect a renamed optional field per record — deliberately delegated to null-rate guard rails | Chosen |
Full strict model, extra="forbid" |
Maximum drift detection; upstream additions surface immediately | Every cosmetic upstream addition rejects the entire population until a redeploy; unacceptable coupling for a free public registry | Rejected |
| Validate the mapped row, not the source document | Cheapest to adopt; catches type and not null violations |
Cannot distinguish "upstream changed" from "our mapper is wrong" — by the time the row exists the mapper has already turned a renamed field into None |
Rejected |
Guard-rail placement¶
| Option | For | Against | Verdict |
|---|---|---|---|
| Inside the load transaction, pre-commit (chosen) | A breach leaves the prior population untouched, reusing the atomicity bulk_load_rows already provides; no window in which a bad population is visible |
Guard-rail queries run against the staging table while holding the swap transaction | Chosen |
| Post-load check, roll forward on breach | Simpler to write | The bad population is live and readable between load and detection; repair requires a second full run | Rejected |
| Separate scheduled audit job | Decoupled; can be arbitrarily expensive | Detects the truncated swap hours after consumers read it; converts a prevention gate into an incident report | Rejected |
Where quality verdicts live¶
| Option | For | Against | Verdict |
|---|---|---|---|
Postgres table + view alongside ingestion_runs (chosen) |
Same store, same transaction, same query surface as the ledger operators already use; joins to the run row for free; no new infrastructure | Grows with run count; needs a retention policy | Chosen |
| Structured logs to Cloud Logging only | No schema work; already wired to alerting | Not joinable to ingestion_runs; log retention is shorter than the questions being asked; cannot express "which gates did not run" |
Rejected |
| External data-quality tool (Great Expectations, Soda) | Mature rule libraries and reporting UI | A second control plane and vocabulary for a solo founder, duplicating the ledger this repo already owns; rule sets would live outside the ADRs that govern them | Rejected for now |
Consequences¶
core/gains a quality layer, parallel to what ADR-0007 did for document fetch: record contracts, the reject ledger, guard rails, the retry policy, and the quality verdict are framework capabilities, not per-source code. Regnskabsdata already proved this framework generalizes across sources; the enrichment pipelines inherit it when they arrive.- Mappers stop taking
dict. They take a validated contract model. This is a mechanical but wide change across both ingesters and their tests, and it removes theKeyError-versus-silent-Nonesplit that currently makes mapper behaviour unpredictable. ingestion_runsalone stops being the answer to "was the run good?" —ingestion_run_reportbecomes the operator's entry point, and the runbook and staleness check point at it. A successful run with a non-empty reject count is a normal, visible state.- Thresholds are a standing calibration obligation. Every threshold in this ADR is an initial estimate, not a measurement — no live full-population run has happened yet for either ingester (the open item recorded against both epics). First production runs must record observed reject rates, null rates, and row-count variance, and this ADR must be amended with them.
- Reject storage carries the ADR-0006 obligation. Quarantining raw
documents is a new persistence site for source data;
deltagerrejects must store the minimized projection. This is a review checkpoint on every future source with personal data, not a one-time fix. - Guard rails can block a legitimate load. A genuine upstream mass deregistration, or a real taxonomy change that legitimately shifts a null rate, will trip a rail and fail the run. That is the intended trade — the override is an explicit, recorded re-run with a widened threshold, never a silent pass.
financial_metricsabsence reasons change consumer semantics. Coverage reporting becomes honest per metric rather than per row, and the read contract's treatment of missing financials can finally be stated precisely.- DBOS remains orthogonal. Nothing here depends on the ADR-0009 spike's outcome, and nothing here is solved by it: DBOS checkpoints execution, these gates decide whether data is fit to commit. A DBOS step that re-runs must re-evaluate the gates, not inherit a prior verdict.
Amendment 2026-09-29: first full people run¶
The first full-register deltager backfill in prod quarantined 8,017 of
1,832,059 documents (0.44%): 8,011 ANDEN_DELTAGER participants with
sidstOpdateret = null, and 3 that also lack enhedstype. The rate is a
stable property of the register, not drift, and a delta never reads these
documents because it ranges on the missing field. The people threshold is
therefore 0.5% (#935). Whether to load them is #936.
Open questions¶
- Every threshold value. Reject rates, and (until the 2026-08-16 amendment above replaced them with one catastrophe threshold) row-count and null-rate bands, and the deletion ceiling are all unmeasured initial estimates. They must be calibrated against the first live full-population runs of CVR and Regnskabsdata, and this ADR amended with the observed distributions.
- Reject and quality-verdict retention. Both tables grow per run; Regnskabsdata's ~6.4M-record population could produce large reject volumes during a genuine upstream incident. Retention and any partitioning are undecided.
- Whether guard-rail evaluation is affordable at full population scale
inside the swap transaction — the null-rate scan over a freshly staged
1.3M-row
companiestable is unmeasured. - Semantic rule coverage for financial coherence. Which cross-metric rules are genuinely sound under Danish ÅRL reporting classes is not settled; ADR-0007 already establishes that small filers legitimately omit turnover and EBIT, so naive coherence rules would fire constantly. Needs the same real-filing sample that G4's residual item needs.
- Whether enrichment pipelines can reuse the reject ledger unchanged or need an evidence-scoped variant — deferred, per this ADR's scope.