Skip to content

Ingestion Quality Gates

The gate reference for registry ingesters, and the operator's answer to "was last night's run good?".

Governed by ADR-0010. Durability guarantees — run ledger, watermark discipline, raw snapshots, staging+swap — are ADR-0005 and are a separate concern: they establish that a run completed and can be replayed, not that what it loaded is correct.

Implementation status

ADR-0010 is accepted. G2 is shipped for every registry entity (#98,

103), G6 records a verdict per run (#99), G4 can fail a run (#100), and

G5's load guard rails are shipped (#101), G1 classifies retries (#102), and G3's semantic rules are shipped (#104). Every gate in the sequence below is now enforced.

What remains is calibration: every threshold here is an unmeasured estimate until #106 replaces it with a number from a real full-population run.

A verdict with a null threshold is deliberate: the gate reports what it measured without judging it. Every threshold in ADR-0010 is an unmeasured estimate until #106 calibrates one against a full-population run. Delivery is tracked by PRD #97 and the slices under Delivery status. Every gate slice is shipped; only threshold calibration (#106) remains open.

Invalid XML syntax, forbidden XML entities, and invalid XBRL context dates are document parse failures. The metrics row records parse_failed, and processing continues with the next document. Unexpected parser errors fail the run. Registry scroll cleanup also runs when an allocated scroll returns malformed data.

Retry attempts share one monotonic completion deadline. Each transport receives the remaining time and limits its connect/read timeout to that value. Response bodies are read in chunks and checked between reads; a late result is rejected before parsing or persistence. This is cooperative enforcement at I/O boundaries: Requests timeouts limit socket inactivity and cannot interrupt every blocked operation at the exact deadline. Owned sessions and responses close on failure.

The gate sequence

Registry run accounting is shared in registry/run_support.py. Registry deltas and participant runs record semantic, transport, record-contract, empty-population, and reject-rate verdicts in that order. The verdicts are flushed before success is committed. An empty delta, or one with only older timestamps, holds the previous watermark. Population swaps check the reject rate before publication.

All three run paths roll back before they record and flush a failure verdict. This keeps the failure evidence after the failed load is discarded. Semantic rejects use the first broken rule's code and join all broken rule descriptions with ;. Each rejected document counts once and contributes no watermark. Participant data is minimized before it reaches this shared recording step.

Every record and every run passes through the same ordered gates. Each gate declares where it runs, what it does to a failing record, and what it writes to the run's quality verdict.

flowchart TD
    A[Upstream scroll / document fetch] --> B{G1 · Transport}
    B -->|terminal 4xx| F1[fail run]
    B -->|retryable, exhausted| F1
    B -->|ok| C{G2 · Record contract}
    C -->|invalid| Q[ingestion_rejects]
    C -->|valid| D{G3 · Semantic rules}
    D -->|severity reject| Q
    D -->|severity warn| E
    D -->|ok| E{G4 · Reject-rate threshold}
    Q --> E
    E -->|breach| F2[fail run · hold watermark]
    E -->|under threshold| G{G5 · Load guard rails}
    G -->|breach| F3[abort swap · population unchanged]
    G -->|ok| H[commit · advance watermark]
    H --> I[G6 · Quality verdict written]
    F1 --> I
    F2 --> I
    F3 --> I
# Gate Runs at Failing record Failing run
G1 Transport (shipped) — retryable/terminal classification, exponential backoff with jitter, Retry-After adapter n/a terminal 4xx, or retries exhausted
G2 Record contract (shipped) — typed structural validation of the source document before mapping quarantined only via G4
G3 Semantic rules (shipped) — value plausibility after mapping quarantined (reject) or counted (warn) only via G4
G4 Reject-rate threshold (shipped) — rejects ÷ documents read end of record loop, pre-commit n/a breach → rollback, watermark held
G5 Load guard rails (shipped) — row-count delta, null-rate delta, deletion ceiling inside the load transaction n/a breach → swap aborted, prior population intact
G6 Quality verdict (shipped) — per-gate outcome recorded end of every run, including failed ones n/a n/a
— Freshness (shipped) — staleness check separate scheduled job n/a breach → non-zero exit → job-failure alert
— Empty-population guard (shipped) — refuse zero-row swap inside the load transaction n/a G5's degenerate case

Null-rate scan cost, measured

ADR-0010 flagged the scan's affordability inside the swap transaction as unverified. Measured on 2026-08-15: 1.44s for 2,260,000 rows across 27 promoted columns, one pass per table. The swap already holds the table for the length of its COPY, which is minutes at that scale, so the rail is not what makes a swap slow.

Verifying a run

ingestion_run_report is the entry point. One query per run answers what ran, what it read, what it loaded, what it rejected and why, and which gates held.

-- Last night, everywhere.
select source, entity, run_type, status,
       docs_read, rows_upserted, rejects, gates_failed
from ingestion_run_report
where started_at > now() - interval '24 hours'
order by started_at desc;

A run with status = 'succeeded' and a non-zero rejects count is a normal, expected state — that is the point of quarantine. What warrants attention is a reject count that moves.

-- Which gates were evaluated on one run, and what did they measure?
select gate, outcome, measured_value, threshold, contract_version
from ingestion_run_quality
where run_id = :run_id
order by gate;

outcome = 'skipped' is meaningful, not neutral: it means a gate did not run (no baseline for a delta rail on a first run, for example). A gate missing from a run's verdict entirely is a defect, not a pass.

-- What exactly was rejected, and why?
select stage, reason_code, count(*) as n
from ingestion_rejects
where run_id = :run_id
group by stage, reason_code
order by n desc;

-- Then inspect individual documents.
select source_key, stage, reason_code, reason_detail, raw
from ingestion_rejects
where run_id = :run_id and reason_code = :code
limit 20;

Rejected deltager documents store the minimized projection only (ADR-0006) — the same allowlist minimize_deltager_for_extract applies to the GCS extract. Do not add a source with personal data to the reject ledger without re-checking this.

Investigating a blocked run

Reject-rate breach (G4)

The run failed, the watermark held, and the next scheduled run will re-read the same window. Nothing was loaded.

  1. Group the run's rejects by reason_code (query above). A single dominant code across most records means upstream shape drift, not bad data.
  2. Compare contract_version in the verdict against the current code. If the contract changed recently, the breach may be ours.
  3. If upstream genuinely changed, update the record contract and ship it — do not widen the threshold to get past it.
  4. If upstream sent a genuine burst of malformed records that we intend to tolerate, widen the entity's threshold as a recorded decision and amend ADR-0010 with the observed rate.

Guard-rail breach (G5)

The swap was aborted. The previous population is intact and readable — this is the failure mode the gate exists to produce.

  1. Read the failing rail's measured_value and threshold from the verdict.
  2. Row-count shortfall almost always means a truncated scroll (upstream partial outage, expired scroll cursor), not a real population change. Re-run before doing anything else.
  3. Null-rate jump on one column means an upstream field rename or restructure. Find the column, check the raw NDJSON for that run in GCS, fix the mapper and contract.
  4. Deletion-ceiling breach means a run tried to remove more rows than allowed. Verify against the source before overriding.

A legitimate upstream mass change (a real deregistration wave, a taxonomy revision that genuinely shifts a null rate) will trip a rail. The override is an explicit re-run with a widened, recorded threshold — never a silent pass.

Absence reasons

financial_metric_coverage records why a promoted metric is NULL — one row per publication and promoted metric — so an operator can distinguish "this company does not file turnover" from "our extractor could not read it" (#105).

Reason Meaning
no_xbrl_document Publication is PDF/paper-only (~37% of filings, ADR-0007); nothing was parsed
not_parsed_yet An XBRL document exists, but no metrics row has been written for it yet
fact_absent Filing read, concept not present — legitimate for small ÅRL reporting classes, which may omit turnover and EBIT
ambiguous_context More than one whole-enterprise value for this concept and period; not guessed at
parse_failed Document fetched but could not be parsed
reason_not_recorded Row parsed before reasons existed; reasons are not backfilled

A publication outside the five-year window (ADR-0007) has no financial_reports row at all, so it is absent from the view rather than carrying a reason.

Only ambiguous_context and parse_failed indicate an extractor problem. The rest are honest coverage limits and are expected at scale.

Consumers filtering on financials must treat every absence as "unknown", never as zero — ADR-0007's rule, now enforceable rather than only documented.

Thresholds

Every threshold is currently an estimate

No live full-population run has happened for either ingester. All values below are initial estimates, not measurements. Calibrating them against the first production runs is an explicit open item in ADR-0010 — until then, expect false positives, and record what you observe.

Reject-rate thresholds are declared config per source/entity, so calibration does not require a code change to the framework.

G5's rails are not. Since 2026-08-16 (#251) all three share one constant, CATASTROPHIC_SHARE, and nothing about them is per-entity: they no longer ask whether a population is different — a question with legitimate answers that refused a real swap the first time they ran — but whether it is broken.

Gate Parameter Value
G4 Reject rate, CVR entities 0.1% (estimate)
G4 Reject rate, Regnskabsdata 1% (estimate)
G5 Row count: share of the population lost 99.5%
G5 Per-column null rate, staged, when the live column was populated 99.5%
G5 Deletion ceiling: share of rows one run may remove 99.5%
— Staleness (shipped) 30h — 1-day contract + 6h margin

What these gates do not cover

  • Execution durability. Resuming an interrupted process is ADR-0009's DBOS spike. A DBOS step that re-runs must re-evaluate these gates; it never inherits a prior verdict.
  • Enrichment pipelines. ADR-0010 scopes registry ingesters only. Enrichment inherits the framework; its own gates (evidence quality, budget enforcement, LLM output validation) are a later decision.
  • Read-contract correctness. View semantics and consumer guarantees are read-contract.md.

Delivery status

Tracked by PRD #97. Slices are listed in dependency order.

Slice Gate Issue
R0 Record contracts + reject ledger (CVR companies tracer) — shipped #98
R1 Per-run quality verdict + ingestion_run_report — shipped #99
R2 Reject-rate threshold (G4) — shipped #100
R3 Load guard rails (G5) — shipped #101
R4 Retry classification (G1) — shipped #102
R5 Contracts for remaining entities — shipped #103
R6 Semantic rules (G3) — shipped #104
R7 Absence reasons — shipped #105
R8 Threshold calibration from live runs #106

Web enrichment (#153)

The registry family's machinery (record contracts, reject ledger, ingestion_run_report) is shipped through R7 (#98–#105); only threshold calibration (#106) remains open. Web enrichment has its own verdict, written in ADR-0010's vocabulary so the two stay aligned rather than diverge.

web_campaign_verdict holds one row per gate per run: gate name, outcome (pass/warn/fail/skipped), measured value, threshold, detail, the rule-set version in force, and the enforcement mode that produced it. It is distinct from web_run_summary (#150), which answers only what the run did: durability and quality are separate guarantees with separate records.

The gates

Gate Measures Threshold
unresolved-rate companies that resolved to nothing 0.60
blocked-rate companies blocked (challenge, 401/403/429, blocked-review) 0.25
failed-rate companies that hit a stage error 0.05
budget-exhausted-rate companies that ran out of budget 0.20
stale-presence-rate published signals older than the freshness tolerance 0.50

failed-rate is held an order of magnitude tighter than the others on purpose: an unresolved company is a fact about the web, and a failed one is a defect in the pipeline.

A gate that cannot be evaluated -- no companies in the run, nothing published yet -- is recorded as skipped with a reason. Its absence is a finding, not a silent pass.

Every threshold is an estimate

No live campaign has run (#196), so these numbers come from what the sibling PoC saw over a few hundred companies, not from this pipeline's own distribution. Enforcement is therefore warn: a breach records a verdict, emits dbos_web_quality_gate_breached for an alert to read, and does not fail the job. Turning that up after calibration is one setting (Enforcement), not a rewrite -- and the recorded rows keep saying which mode produced them, so an old verdict stays readable as the estimate it was.

This is the same problem #106 records for the registry family, with the same fix: measure a live run, then calibrate.

The freshness contract

A resolved website stays valid for 30 days -- deliberately the same number as the page cache's time to live (#151), so the contract and the mechanism that enforces it cannot drift apart. Re-enrichment is triggered by running a campaign over the company again; the cache TTL then decides whether that re-run pays for the page or reuses it. There is no timer: the next campaign is the rebuild path. stale-presence-rate is what makes the tolerance measurable rather than merely documented.

Not built: precision per extractor version

153 also asks for a precision measure against a labelled sample. No

labelled sample exists, and precision without labels is not measurable, so it is filed separately as #212 rather than approximated. Confidence values today (0.98 CVR-grep, 0.9 site-declared social, 0.6 SERP) remain fixed constants calibrated against nothing. The verdict counts how many companies resolved; it says nothing about how many resolved correctly.