Ingestion Quality Gates¶
The gate reference for registry ingesters, and the operator's answer to "was last night's run good?".
Governed by ADR-0010. Durability guarantees — run ledger, watermark discipline, raw snapshots, staging+swap — are ADR-0005 and are a separate concern: they establish that a run completed and can be replayed, not that what it loaded is correct.
Implementation status
ADR-0010 is accepted. G2 is shipped for every registry entity (#98,
103), G6 records a verdict per run (#99), G4 can fail a run (#100), and¶
G5's load guard rails are shipped (#101), G1 classifies retries (#102), and G3's semantic rules are shipped (#104). Every gate in the sequence below is now enforced.
What remains is calibration: every threshold here is an unmeasured estimate until #106 replaces it with a number from a real full-population run.
A verdict with a null threshold is deliberate: the gate reports what it
measured without judging it. Every threshold in ADR-0010 is an unmeasured
estimate until #106 calibrates one against a full-population run.
Delivery is tracked by PRD #97
and the slices under Delivery status. Every gate slice
is shipped; only threshold calibration (#106) remains open.
Invalid XML syntax, forbidden XML entities, and invalid XBRL context dates are
document parse failures. The metrics row records parse_failed, and processing
continues with the next document. Unexpected parser errors fail the run.
Registry scroll cleanup also runs when an allocated scroll returns malformed data.
Retry attempts share one monotonic completion deadline. Each transport receives the remaining time and limits its connect/read timeout to that value. Response bodies are read in chunks and checked between reads; a late result is rejected before parsing or persistence. This is cooperative enforcement at I/O boundaries: Requests timeouts limit socket inactivity and cannot interrupt every blocked operation at the exact deadline. Owned sessions and responses close on failure.
The gate sequence¶
Registry run accounting is shared in registry/run_support.py. Registry deltas
and participant runs record semantic, transport, record-contract, empty-population,
and reject-rate verdicts in that order. The verdicts are flushed before success is
committed. An empty delta, or one with only older timestamps, holds the previous
watermark. Population swaps check the reject rate before publication.
All three run paths roll back before they record and flush a failure verdict.
This keeps the failure evidence after the failed load is discarded. Semantic
rejects use the first broken rule's code and join all broken rule descriptions
with ;. Each rejected document counts once and contributes no watermark.
Participant data is minimized before it reaches this shared recording step.
Every record and every run passes through the same ordered gates. Each gate declares where it runs, what it does to a failing record, and what it writes to the run's quality verdict.
flowchart TD
A[Upstream scroll / document fetch] --> B{G1 · Transport}
B -->|terminal 4xx| F1[fail run]
B -->|retryable, exhausted| F1
B -->|ok| C{G2 · Record contract}
C -->|invalid| Q[ingestion_rejects]
C -->|valid| D{G3 · Semantic rules}
D -->|severity reject| Q
D -->|severity warn| E
D -->|ok| E{G4 · Reject-rate threshold}
Q --> E
E -->|breach| F2[fail run · hold watermark]
E -->|under threshold| G{G5 · Load guard rails}
G -->|breach| F3[abort swap · population unchanged]
G -->|ok| H[commit · advance watermark]
H --> I[G6 · Quality verdict written]
F1 --> I
F2 --> I
F3 --> I
| # | Gate | Runs at | Failing record | Failing run |
|---|---|---|---|---|
| G1 | Transport (shipped) — retryable/terminal classification, exponential backoff with jitter, Retry-After |
adapter | n/a | terminal 4xx, or retries exhausted |
| G2 | Record contract (shipped) — typed structural validation of the source document | before mapping | quarantined | only via G4 |
| G3 | Semantic rules (shipped) — value plausibility | after mapping | quarantined (reject) or counted (warn) |
only via G4 |
| G4 | Reject-rate threshold (shipped) — rejects ÷ documents read | end of record loop, pre-commit | n/a | breach → rollback, watermark held |
| G5 | Load guard rails (shipped) — row-count delta, null-rate delta, deletion ceiling | inside the load transaction | n/a | breach → swap aborted, prior population intact |
| G6 | Quality verdict (shipped) — per-gate outcome recorded | end of every run, including failed ones | n/a | n/a |
| — | Freshness (shipped) — staleness check | separate scheduled job | n/a | breach → non-zero exit → job-failure alert |
| — | Empty-population guard (shipped) — refuse zero-row swap | inside the load transaction | n/a | G5's degenerate case |
Null-rate scan cost, measured
ADR-0010 flagged the scan's affordability inside the swap transaction as unverified. Measured on 2026-08-15: 1.44s for 2,260,000 rows across 27 promoted columns, one pass per table. The swap already holds the table for the length of its COPY, which is minutes at that scale, so the rail is not what makes a swap slow.
Verifying a run¶
ingestion_run_report is the entry point. One query per run answers what ran,
what it read, what it loaded, what it rejected and why, and which gates held.
-- Last night, everywhere.
select source, entity, run_type, status,
docs_read, rows_upserted, rejects, gates_failed
from ingestion_run_report
where started_at > now() - interval '24 hours'
order by started_at desc;
A run with status = 'succeeded' and a non-zero rejects count is a normal,
expected state — that is the point of quarantine. What warrants attention is
a reject count that moves.
-- Which gates were evaluated on one run, and what did they measure?
select gate, outcome, measured_value, threshold, contract_version
from ingestion_run_quality
where run_id = :run_id
order by gate;
outcome = 'skipped' is meaningful, not neutral: it means a gate did not run
(no baseline for a delta rail on a first run, for example). A gate missing
from a run's verdict entirely is a defect, not a pass.
-- What exactly was rejected, and why?
select stage, reason_code, count(*) as n
from ingestion_rejects
where run_id = :run_id
group by stage, reason_code
order by n desc;
-- Then inspect individual documents.
select source_key, stage, reason_code, reason_detail, raw
from ingestion_rejects
where run_id = :run_id and reason_code = :code
limit 20;
Rejected deltager documents store the minimized projection only
(ADR-0006) — the same
allowlist minimize_deltager_for_extract applies to the GCS extract. Do not
add a source with personal data to the reject ledger without re-checking this.
Investigating a blocked run¶
Reject-rate breach (G4)¶
The run failed, the watermark held, and the next scheduled run will re-read the same window. Nothing was loaded.
- Group the run's rejects by
reason_code(query above). A single dominant code across most records means upstream shape drift, not bad data. - Compare
contract_versionin the verdict against the current code. If the contract changed recently, the breach may be ours. - If upstream genuinely changed, update the record contract and ship it — do not widen the threshold to get past it.
- If upstream sent a genuine burst of malformed records that we intend to tolerate, widen the entity's threshold as a recorded decision and amend ADR-0010 with the observed rate.
Guard-rail breach (G5)¶
The swap was aborted. The previous population is intact and readable — this is the failure mode the gate exists to produce.
- Read the failing rail's
measured_valueandthresholdfrom the verdict. - Row-count shortfall almost always means a truncated scroll (upstream partial outage, expired scroll cursor), not a real population change. Re-run before doing anything else.
- Null-rate jump on one column means an upstream field rename or restructure. Find the column, check the raw NDJSON for that run in GCS, fix the mapper and contract.
- Deletion-ceiling breach means a run tried to remove more rows than allowed. Verify against the source before overriding.
A legitimate upstream mass change (a real deregistration wave, a taxonomy revision that genuinely shifts a null rate) will trip a rail. The override is an explicit re-run with a widened, recorded threshold — never a silent pass.
Absence reasons¶
financial_metric_coverage records why a promoted metric is NULL — one row
per publication and promoted metric — so an operator can distinguish "this
company does not file turnover" from "our extractor could not read it" (#105).
| Reason | Meaning |
|---|---|
no_xbrl_document |
Publication is PDF/paper-only (~37% of filings, ADR-0007); nothing was parsed |
not_parsed_yet |
An XBRL document exists, but no metrics row has been written for it yet |
fact_absent |
Filing read, concept not present — legitimate for small ÅRL reporting classes, which may omit turnover and EBIT |
ambiguous_context |
More than one whole-enterprise value for this concept and period; not guessed at |
parse_failed |
Document fetched but could not be parsed |
reason_not_recorded |
Row parsed before reasons existed; reasons are not backfilled |
A publication outside the five-year window (ADR-0007) has no financial_reports
row at all, so it is absent from the view rather than carrying a reason.
Only ambiguous_context and parse_failed indicate an extractor problem. The
rest are honest coverage limits and are expected at scale.
Consumers filtering on financials must treat every absence as "unknown", never as zero — ADR-0007's rule, now enforceable rather than only documented.
Thresholds¶
Every threshold is currently an estimate
No live full-population run has happened for either ingester. All values below are initial estimates, not measurements. Calibrating them against the first production runs is an explicit open item in ADR-0010 — until then, expect false positives, and record what you observe.
Reject-rate thresholds are declared config per source/entity, so calibration does not require a code change to the framework.
G5's rails are not. Since 2026-08-16 (#251) all three share one constant,
CATASTROPHIC_SHARE, and nothing about them is per-entity: they no longer ask
whether a population is different — a question with legitimate answers that
refused a real swap the first time they ran — but whether it is broken.
| Gate | Parameter | Value |
|---|---|---|
| G4 | Reject rate, CVR entities | 0.1% (estimate) |
| G4 | Reject rate, Regnskabsdata | 1% (estimate) |
| G5 | Row count: share of the population lost | 99.5% |
| G5 | Per-column null rate, staged, when the live column was populated | 99.5% |
| G5 | Deletion ceiling: share of rows one run may remove | 99.5% |
| — | Staleness (shipped) | 30h — 1-day contract + 6h margin |
What these gates do not cover¶
- Execution durability. Resuming an interrupted process is ADR-0009's DBOS spike. A DBOS step that re-runs must re-evaluate these gates; it never inherits a prior verdict.
- Enrichment pipelines. ADR-0010 scopes registry ingesters only. Enrichment inherits the framework; its own gates (evidence quality, budget enforcement, LLM output validation) are a later decision.
- Read-contract correctness. View semantics and consumer guarantees are read-contract.md.
Delivery status¶
Tracked by PRD #97. Slices are listed in dependency order.
| Slice | Gate | Issue |
|---|---|---|
| R0 | Record contracts + reject ledger (CVR companies tracer) — shipped | #98 |
| R1 | Per-run quality verdict + ingestion_run_report — shipped |
#99 |
| R2 | Reject-rate threshold (G4) — shipped | #100 |
| R3 | Load guard rails (G5) — shipped | #101 |
| R4 | Retry classification (G1) — shipped | #102 |
| R5 | Contracts for remaining entities — shipped | #103 |
| R6 | Semantic rules (G3) — shipped | #104 |
| R7 | Absence reasons — shipped | #105 |
| R8 | Threshold calibration from live runs | #106 |
Web enrichment (#153)¶
The registry family's machinery (record contracts, reject ledger,
ingestion_run_report) is shipped through R7 (#98–#105); only threshold
calibration (#106) remains open. Web enrichment has its own verdict,
written in ADR-0010's vocabulary so the two stay aligned rather than
diverge.
web_campaign_verdict holds one row per gate per run: gate name, outcome
(pass/warn/fail/skipped), measured value, threshold, detail, the
rule-set version in force, and the enforcement mode that produced it. It is
distinct from web_run_summary (#150), which answers only what the run
did: durability and quality are separate guarantees with separate records.
The gates¶
| Gate | Measures | Threshold |
|---|---|---|
unresolved-rate |
companies that resolved to nothing | 0.60 |
blocked-rate |
companies blocked (challenge, 401/403/429, blocked-review) | 0.25 |
failed-rate |
companies that hit a stage error | 0.05 |
budget-exhausted-rate |
companies that ran out of budget | 0.20 |
stale-presence-rate |
published signals older than the freshness tolerance | 0.50 |
failed-rate is held an order of magnitude tighter than the others on
purpose: an unresolved company is a fact about the web, and a failed one is
a defect in the pipeline.
A gate that cannot be evaluated -- no companies in the run, nothing
published yet -- is recorded as skipped with a reason. Its absence is a
finding, not a silent pass.
Every threshold is an estimate¶
No live campaign has run (#196), so these numbers come from what the
sibling PoC saw over a few hundred companies, not from this pipeline's own
distribution. Enforcement is therefore warn: a breach records a verdict,
emits dbos_web_quality_gate_breached for an alert to read, and does not
fail the job. Turning that up after calibration is one setting
(Enforcement), not a rewrite -- and the recorded rows keep saying which
mode produced them, so an old verdict stays readable as the estimate it was.
This is the same problem #106 records for the registry family, with the same fix: measure a live run, then calibrate.
The freshness contract¶
A resolved website stays valid for 30 days -- deliberately the same
number as the page cache's time to live (#151), so the contract and the
mechanism that enforces it cannot drift apart. Re-enrichment is triggered
by running a campaign over the company again; the cache TTL then decides
whether that re-run pays for the page or reuses it. There is no timer: the
next campaign is the rebuild path. stale-presence-rate is what makes the
tolerance measurable rather than merely documented.
Not built: precision per extractor version¶
153 also asks for a precision measure against a labelled sample. No¶
labelled sample exists, and precision without labels is not measurable, so it is filed separately as #212 rather than approximated. Confidence values today (0.98 CVR-grep, 0.9 site-declared social, 0.6 SERP) remain fixed constants calibrated against nothing. The verdict counts how many companies resolved; it says nothing about how many resolved correctly.