Skip to content

Regnskabsdata ingestion: same-host ES surface, eager nøgletal on delta, bounded five-year backfill

Status: accepted Date: 2026-07-10 Deciders: Chris (solo founder) Depends on: ADR-0011 (local hub and ownership boundary) and ADR-0005 (ingestion architecture, watermark discipline, raw-store rule) Supersedes (partially): prephase doc 04's lazy-document model and its indlaesningsId idempotency key Research: related-work synthesis, gap analysis

Context

Regnskabsdata (digitally filed annual reports) is the second ingester and the first reuse of the framework the CVR ingester establishes (adapter/ingester split, core/ loaders, run ledger, raw store). It feeds financial_reports (filing metadata, 1:N to companies) and financial_metrics (promoted nøgletal used for population-wide financial filtering).

Live probes (2026-07-10, recorded in the research docs) established:

  • http://distribution.virk.dk/offentliggoerelser/_search is free and credential-less (alias for index indberetninger-20230623, same host as cvr-permanent). Plain HTTP, like the CVR surface.
  • 6,410,977 records, 100% offentliggoerelsestype: regnskab — no population filter needed.
  • sidstOpdateret, offentliggoerelsesTidspunkt, indlaesningsTidspunkt, and the accounting period are present on 100% of records; indlaesningsId is missing on 54% (3,440,462) — unusable as a key. The ES _id (urn:ofk:oid:…) is present on every hit.
  • 4,028,669 records (63%) carry an XBRL (XML) document; the rest are PDF/paper-only. 19,727 records (~0.3%) are corrections (omgoerelse).
  • Delta volume ≈ 15.7k updates/7 days off-season; filing season (May–June) bursts higher.
  • The cluster self-reports Elasticsearch 6.8.23 — contradicting the "ES 1.7 dialect" assumption recorded for the CVR surface (see amendment note in ADR-0005's open questions).

The tension forcing this ADR: prephase doc 04 wanted the XBRL-derived nøgletal both filterable (which demands population-wide coverage — search indexes cannot be lazy-loaded, per prephase ADR-0004) and lazily fetched on drill-in (which yields coverage only where a user already looked). Both cannot hold. Additionally, this repository's service boundary gives consumer repos no path to trigger a lazy fetch: they read the hub and never call adapters.

Decision

  1. Same surface family, no credential. The regnskab_es adapter targets distribution.virk.dk/offentliggoerelser — same host, dialect, and HTTP-only constraint as the CVR adapter (ingestion runs only from Cloud Run, per ADR-0005), minus Basic Auth. Document downloads (regnskaber.virk.dk) inherit the same constraint.
  2. Metadata is population-wide and incremental, exactly like CVR. Full-scroll backfill of all 6.41M publication records via staging+swap, then daily delta on sidstOpdateret ≥ watermark with transactional upserts, monthly reconciliation, every run stamped in ingestion_runs. The watermark discipline of ADR-0005 transfers unchanged.
  3. Nøgletal are fetched eagerly, not lazily.
  4. The daily delta run fetches and parses the XBRL document inline for each new/updated publication that has one, and upserts financial_metrics in the same run. At observed volumes this is a few hundred to a few thousand documents per day.
  5. History is covered by a bounded backfill: the last five filing years, newest first (~1.5–2M documents), run as a resumable batch job tracked in ingestion_runs. Older periods stay metadata-only.
  6. Widening the window later is a re-run with an earlier cutoff — cheap by construction, so deferring deep history costs nothing structurally.
  7. Corrections append; reads use a latest-view. Every publication is its own row, keyed on load_id = the ES _id (present on 100% of records; indlaesningsId is dead). There is no unique (cvr, period_start, period_end) constraint; is_correction (omgoerelse) is stored as fact. The read contract exposes a financial_reports_latest view — newest offentliggoerelsesTidspunkt per (cvr, accounting period). Filing history is provenance, not an update anomaly.
  8. Raw XBRL lives in GCS; Postgres gets promoted columns only. The fetched XBRL file lands in GCS under the run-scoped raw-store convention; the extractor promotes the ~10 filterable nøgletal into financial_metrics columns. No xbrl_facts JSONB in v1 — a future need re-parses from GCS (extractor is versioned and re-runnable), never re-fetches upstream. Publication metadata rows keep a raw JSONB column like companies (the metadata doc is small); no personal data is involved (ADR-0006 does not bite here).

Options considered

Nøgletal population model

Option For Against Verdict
Eager on delta + bounded 5-year backfill (chosen) Population-wide filterability for the periods products show; no cross-service fetch trigger invented; reuses bulk machinery; window widening is a cheap re-run Historical metrics beyond 5 years absent until someone widens the window Chosen
Fully eager, all-time (~4M docs) Complete corpus Weeks of polite-rate fetching and storage for pre-2021 periods no scenario uses; oldest taxonomies are the hardest to parse and least valuable Rejected
Strictly lazy (prephase doc 04) Minimal fetching Breaks filterability (coverage only where users looked); requires a consumer→hub fetch-request mechanism that violates the service boundary Rejected
Shortlist-scoped enrichment pipeline Fits existing enrichment vocabulary Makes financial filtering permanently shortlist-bound; Regnskabsdata is a class-A bulk source, not paid enrichment Rejected

Corrections

Option For Against Verdict
Append + latest-view (chosen) Trivial idempotency (upsert on load_id); history preserved; correction logic lives in one view Table carries ~0.3% superseded rows Chosen
Supersede in place One row per (cvr, period) Destroys filing history; replay ambiguity; breaks on the first legitimate re-filing Rejected
Append + superseded_by pointer Explicit lineage Loader-side ordering logic and second writes for a case the view answers Rejected

XBRL storage

Option For Against Verdict
GCS file + promoted columns (chosen) Raw-snapshot reproducibility; bounded Postgres size; versioned extractor re-runs from GCS Ad-hoc fact queries need a re-parse pass Chosen
+ xbrl_facts JSONB in Postgres Ad-hoc querying now Millions of large taxonomy-versioned JSONB blobs before any consumer asks Rejected
Postgres promoted columns only, discard document Cheapest Extractor bug ⇒ re-fetch millions of docs upstream; violates the raw-snapshot rule Rejected

Consequences

  • core/ gains a document-fetch stage (politeness-budgeted HTTP fetch + raw-store write + parse hook) — the template every later document-heavy source (Tinglysning, EMO) reuses. It does not gain lazy-read machinery; no details_cached_at semantics ship in v1.
  • financial_metrics coverage is honest and bounded: ~63% of reports have XBRL at all, and only the 5-year window is parsed. Coverage per company is reportable (missing ≠ zero); consumers filtering on financials must treat absent metrics as "not filed digitally / outside window", never as a value.
  • The XBRL→nøgletal extractor is a versioned contract: Danish taxonomies change yearly, so extraction is per-taxonomy-version, and an extractor bump triggers re-parse from GCS, stamped by extractor version.
  • Filing-season bursts make delta run duration elastic (thousands of documents on peak days); the run must be resumable and rate-polite rather than time-boxed.
  • The delta and the metrics fetch are one run: a document-fetch failure fails the run without advancing the watermark (ADR-0005 discipline), so metrics can lag metadata only within a run, never silently across runs.

Open questions

  • Real fetch tolerance and document sizes on regnskaber.virk.dk (measure in the tracer bullet; G6).
  • G4 — updated 2026-07-19, not fully closed (issue #54, taxonomy spike): promoted set is assets, equity, result, gross_profit, revenue, ebit, ebitda (candidate fsa: element names per column in the spike note); extraction approach decided as a narrow hand-rolled extractor, validated offline per taxonomy version against an Arelle-resolved sample rather than run as a runtime dependency. revenue/ebit/ebitda are expected to be structurally low-coverage (Danish ÅRL lets small filers omit them, not a parsing bug). Residual open item: none of this is verified against a real filing yet — run the spike's ~20-filing sample before the financial_metrics migration is written, per the note.
  • Whether sidstOpdateret moves on correction re-publication (G7 residual; verify during the tracer bullet — the reconciliation path covers misses either way).

Status update (2026-07-20, issue #58 — epic closure)

All five R0–R4 slices (#54–#58) shipped. Every genuinely new framework capability this ADR called for landed: the no-credential adapter variant, the load_id/ES-_id key (G5, closed), the document-fetch capability (core/doc_fetch.py), the versioned XBRL extractor, the bounded backfill, and reconciliation + read contract reusing (and, for _run's pre_swap_hook, further generalizing) the CVR framework unchanged. The financial-band read-contract view intentionally covers a single-period band filter, not PROJECT_CHARTER.md's literal multi-year "positive gross-profit trend" — recorded as a residual scope gap in docs/reference/read-contract.md, not silently claimed as done.

What remains genuinely open, unresolved by this epic: every item above this note that depends on a live filing or a live Cloud Run execution — regnskaber.virk.dk fetch tolerance/sizes (G6), the taxonomy element-name table's verification against ~20 real filings (G4 residual), whether sidstOpdateret moves on correction re-publication (G7 residual), and the first real end-to-end run of delta + eager-metrics + reconcile against production traffic. None of these were fabricated as "measured" or "verified" anywhere in this epic — every fixture and taxonomy claim carries an explicit constructed/unverified flag, consistently, across #54–#58.

Amendment 2026-08-12 — Inline XBRL and whole-enterprise facts (#176, #177)

The extractor remains narrow and hand-rolled, but its input contract includes Inline XBRL 2008 and 2013. Its fact QName is read from name; numeric scale, sign, and the observed Danish ixt:numdotcomma transformation are applied before promotion. A candidate context must have the report period end and no segment or scenario, because either location makes it a dimensional breakdown rather than a whole-enterprise total. Repeated facts with the same promoted value are one observation; conflicting values remain ambiguous and are left null. This is extractor version v2, so affected raw snapshots must be reparsed by the bounded metrics backfill.

Amendment 2026-08-12 — the statement, not seven figures, and derived vs disclosed (#188)

Decision #3 promoted a small filterable set of nøgletal on the reasoning that everything else stays in the raw snapshot. Measured against a commercial product, that set was too small to be competitive: comparing our output for LASSO X A/S (CVR 34580820) against the income statement, balance sheet and nøgletal proff.dk publishes for the same company showed every figure they display is already inside the XBRL document this pipeline fetches, parses and stores. We were promoting 6 usable facts out of roughly 45. The cost of promoting the rest is a wider column list over data already in hand; the cost of not promoting them is that every consumer re-parses the raw snapshot to ask an ordinary question.

The promoted set therefore widens to the statement itself: the balance sheet (liabilities_and_equity, contributed_capital, retained_earnings, noncurrent_assets, current_assets, cash, trade_receivables, provisions, shortterm_liabilities, longterm_liabilities), the income statement (staff_costs, depreciation, profit_before_tax, tax_expense) and employees — an exact headcount where companies.employees_band only ever gives a band.

A column holds a fact the company disclosed. Nothing promoted is summed, inferred, or filled from an accounting identity. Two reasons. "Missing != zero" only means something if a null says "the filing did not say this", and a derived value written into the same column is indistinguishable from a disclosed one afterwards. And a figure frozen at parse time goes stale the moment its formula is corrected, while a view recomputes for free. So longterm_liabilities stays null in the 78% of filings that omit it, even though it is recoverable as liabilities_and_equity - equity - provisions - shortterm_liabilities; a consumer that wants the identity can apply it, with the derivation visible.

Consequently ebitda stops being a stored column. It is a pure function of two disclosed facts and is computed by the financial_key_figures view along with soliditetsgrad, likviditetsgrad 1, afkastningsgrad, egenkapitalens forrentning, kapacitetsgrad, dækningsgrad and overskudsgrad. Each formula was verified to reproduce proff.dk's published figures exactly for the filing above. Dækningsgrad and overskudsgrad divide by turnover and are therefore null for most Danish companies — proff.dk leaves them blank for the same reason.

EBITDA's definition follows the market rather than the textbook (#187). The only depreciation concept Danish filings carry bundles impairment losses in with depreciation and amortisation, and proff.dk computes EBITDA from exactly that figure. Being comparable with the products a customer already uses beats being pure about impairment; the deviation is recorded here rather than hidden in a column comment.

This is extractor version v3, so raw snapshots parsed by v1 or v2 must be reparsed by the bounded metrics backfill.

Amendment 2026-08-12 — publication metadata drops raw, and retention is bounded to the metrics window

This amendment changes decision 5 (publication metadata rows keep a raw JSONB column) and decision 2 (metadata is population-wide), and follows ADR-0005's 2026-08-12 amendment, which removes raw from companies and production_units on measured evidence. Decision 3's bounded five-year nøgletal window is unchanged — this amendment aligns metadata retention to that window rather than altering it.

Decision.

  1. Drop the raw column from financial_reports. Decision 5's reasoning — "the metadata doc is small" — is true per document (mean 758 bytes) and false in aggregate, because this is the highest-row-count table in the hub at 6.42M publications. The publication metadata document already lands in GCS: ingesters/regnskab/pipeline.py reuses ingesters.cvr.pipeline's _run/_run_delta unchanged, so the same run-scoped NDJSON extract is written here as for CVR entities. Decision 5's own principle — raw snapshots live in GCS, the extractor is versioned and re-runnable from them — is extended to the metadata document rather than contradicted.
  2. Retain publication metadata only within the five-year window, keyed on period_end >= cutoff — the same predicate and the same column _select_reports_needing_metrics already uses, so the retained set is exactly the set eligible for nøgletal. 1,666,945 of 6,422,033 publications qualify at today's cutoff. This replaces decision 2's "all 6.41M publication records" and narrows decision 4: corrections still append and financial_reports_latest still resolves them, but filing history is preserved only inside the window.

Evidence (measured, same method as the ADR-0005 amendment):

table rows retained bytes/row size
financial_reports with raw, all history 6,422,033 1,092 7.02 GB
financial_reports without raw, all history 6,422,033 170 1.10 GB
financial_reports without raw, 5-year window 1,666,945 170 0.28 GB
financial_metrics (post-#188 columns), 5-year window ~1.05M (63% have XBRL) 336 0.35 GB

Note the honest shape of the second lever: dropping raw is worth 5.9 GB; bounding retention on top of it is worth a further 0.8 GB. The window was agreed as a size measure when the hub was believed to be far larger. It is retained deliberately, for consistency with the metrics window rather than for the bytes.

Consequences.

  • The four-entity hub lands near 2.6 GB (companies 0.49, production_units 0.62, financial_reports 0.28, financial_metrics 0.35, people + roles 0.84). Allow roughly 5 GB in service, for bloat, WAL, and autovacuum lag.
  • Pre-2021 filing metadata leaves Postgres. It stays in the GCS extracts and remains fully re-scrollable from offentliggoerelser, which is credential-free and has no bounded history. Widening the window is still "a re-run with an earlier cutoff", exactly as decision 3 promises — but it is now a re-run of the metadata backfill as well as the metrics one.
  • A retained publication whose accounting period ages out of the window becomes eligible for deletion, and its financial_metrics row goes with it. The FK from financial_metrics.load_id has no cascade, so the child must be deleted first or the delete fails outright — but the deeper point is that keeping derived nøgletal for a period no consumer can filter on would defeat the window. The XBRL document behind each one stays in GCS and the extractor is versioned and re-runnable against it (decision 5), so this is re-derivable, not lost.
  • Retention needs no scheduler of its own. The window is enforced on the scroll — backfill, delta, and reconcile all carry it — so an out-of-window publication is never ingested. The monthly reconcile then removes rows that have since aged out for free: its staging+swap replaces the population with a freshly bounded scroll, and its existing _delete_orphaned_metrics pre-swap hook takes the derived metrics with them. Steady-state retention is therefore at most one reconcile cycle stale, which is immaterial for a window measured in years. The explicit prune command exists for the cutover from the unbounded table and for use between reconciles.
  • financial_reports_latest and the read contract are unaffected in shape. Consumers see fewer rows, not different columns.

Retained replay and bounded processing (#782)

Document acquisition returns the bytes that won the create-only write, their URI, digest, and original acquisition time. Re-extraction reads those bytes; it does not fetch a newer upstream document under an old object identity. Local and GCS document adapters follow the same rule.

Network pacing and max_fetches apply only to remote acquisition. Each run also permits at most 10,000 snapshot reads by default. Exhaustion fails the run after preserving completed metrics; the existing missing-extractor-version selection resumes the next run. Candidate selection uses 250-row keyset pages in the existing period_end desc, load_id desc order. max_docs remains an additional caller limit. No separate progress table is introduced.