Regnskabsdata ingestion: same-host ES surface, eager nøgletal on delta, bounded five-year backfill¶
Status: accepted
Date: 2026-07-10
Deciders: Chris (solo founder)
Depends on: ADR-0011 (local hub and ownership boundary) and ADR-0005 (ingestion architecture, watermark discipline, raw-store rule)
Supersedes (partially): prephase doc 04's lazy-document model and its indlaesningsId idempotency key
Research: related-work synthesis, gap analysis
Context¶
Regnskabsdata (digitally filed annual reports) is the second ingester and the
first reuse of the framework the CVR ingester establishes (adapter/ingester
split, core/ loaders, run ledger, raw store). It feeds financial_reports
(filing metadata, 1:N to companies) and financial_metrics (promoted
nøgletal used for population-wide financial filtering).
Live probes (2026-07-10, recorded in the research docs) established:
http://distribution.virk.dk/offentliggoerelser/_searchis free and credential-less (alias for indexindberetninger-20230623, same host ascvr-permanent). Plain HTTP, like the CVR surface.- 6,410,977 records, 100%
offentliggoerelsestype: regnskab— no population filter needed. sidstOpdateret,offentliggoerelsesTidspunkt,indlaesningsTidspunkt, and the accounting period are present on 100% of records;indlaesningsIdis missing on 54% (3,440,462) — unusable as a key. The ES_id(urn:ofk:oid:…) is present on every hit.- 4,028,669 records (63%) carry an XBRL (XML) document; the rest are
PDF/paper-only. 19,727 records (~0.3%) are corrections (
omgoerelse). - Delta volume ≈ 15.7k updates/7 days off-season; filing season (May–June) bursts higher.
- The cluster self-reports Elasticsearch 6.8.23 — contradicting the "ES 1.7 dialect" assumption recorded for the CVR surface (see amendment note in ADR-0005's open questions).
The tension forcing this ADR: prephase doc 04 wanted the XBRL-derived nøgletal both filterable (which demands population-wide coverage — search indexes cannot be lazy-loaded, per prephase ADR-0004) and lazily fetched on drill-in (which yields coverage only where a user already looked). Both cannot hold. Additionally, this repository's service boundary gives consumer repos no path to trigger a lazy fetch: they read the hub and never call adapters.
Decision¶
- Same surface family, no credential. The
regnskab_esadapter targetsdistribution.virk.dk/offentliggoerelser— same host, dialect, and HTTP-only constraint as the CVR adapter (ingestion runs only from Cloud Run, per ADR-0005), minus Basic Auth. Document downloads (regnskaber.virk.dk) inherit the same constraint. - Metadata is population-wide and incremental, exactly like CVR.
Full-scroll backfill of all 6.41M publication records via staging+swap,
then daily delta on
sidstOpdateret≥ watermark with transactional upserts, monthly reconciliation, every run stamped iningestion_runs. The watermark discipline of ADR-0005 transfers unchanged. - Nøgletal are fetched eagerly, not lazily.
- The daily delta run fetches and parses the XBRL document inline for
each new/updated publication that has one, and upserts
financial_metricsin the same run. At observed volumes this is a few hundred to a few thousand documents per day. - History is covered by a bounded backfill: the last five filing
years, newest first (~1.5–2M documents), run as a resumable batch job
tracked in
ingestion_runs. Older periods stay metadata-only. - Widening the window later is a re-run with an earlier cutoff — cheap by construction, so deferring deep history costs nothing structurally.
- Corrections append; reads use a latest-view. Every publication is its
own row, keyed on
load_id= the ES_id(present on 100% of records;indlaesningsIdis dead). There is nounique (cvr, period_start, period_end)constraint;is_correction(omgoerelse) is stored as fact. The read contract exposes afinancial_reports_latestview — newestoffentliggoerelsesTidspunktper (cvr, accounting period). Filing history is provenance, not an update anomaly. - Raw XBRL lives in GCS; Postgres gets promoted columns only. The
fetched XBRL file lands in GCS under the run-scoped raw-store convention;
the extractor promotes the ~10 filterable nøgletal into
financial_metricscolumns. Noxbrl_factsJSONB in v1 — a future need re-parses from GCS (extractor is versioned and re-runnable), never re-fetches upstream. Publication metadata rows keep arawJSONB column likecompanies(the metadata doc is small); no personal data is involved (ADR-0006 does not bite here).
Options considered¶
Nøgletal population model¶
| Option | For | Against | Verdict |
|---|---|---|---|
| Eager on delta + bounded 5-year backfill (chosen) | Population-wide filterability for the periods products show; no cross-service fetch trigger invented; reuses bulk machinery; window widening is a cheap re-run | Historical metrics beyond 5 years absent until someone widens the window | Chosen |
| Fully eager, all-time (~4M docs) | Complete corpus | Weeks of polite-rate fetching and storage for pre-2021 periods no scenario uses; oldest taxonomies are the hardest to parse and least valuable | Rejected |
| Strictly lazy (prephase doc 04) | Minimal fetching | Breaks filterability (coverage only where users looked); requires a consumer→hub fetch-request mechanism that violates the service boundary | Rejected |
| Shortlist-scoped enrichment pipeline | Fits existing enrichment vocabulary | Makes financial filtering permanently shortlist-bound; Regnskabsdata is a class-A bulk source, not paid enrichment | Rejected |
Corrections¶
| Option | For | Against | Verdict |
|---|---|---|---|
| Append + latest-view (chosen) | Trivial idempotency (upsert on load_id); history preserved; correction logic lives in one view |
Table carries ~0.3% superseded rows | Chosen |
| Supersede in place | One row per (cvr, period) | Destroys filing history; replay ambiguity; breaks on the first legitimate re-filing | Rejected |
Append + superseded_by pointer |
Explicit lineage | Loader-side ordering logic and second writes for a case the view answers | Rejected |
XBRL storage¶
| Option | For | Against | Verdict |
|---|---|---|---|
| GCS file + promoted columns (chosen) | Raw-snapshot reproducibility; bounded Postgres size; versioned extractor re-runs from GCS | Ad-hoc fact queries need a re-parse pass | Chosen |
+ xbrl_facts JSONB in Postgres |
Ad-hoc querying now | Millions of large taxonomy-versioned JSONB blobs before any consumer asks | Rejected |
| Postgres promoted columns only, discard document | Cheapest | Extractor bug ⇒ re-fetch millions of docs upstream; violates the raw-snapshot rule | Rejected |
Consequences¶
core/gains a document-fetch stage (politeness-budgeted HTTP fetch + raw-store write + parse hook) — the template every later document-heavy source (Tinglysning, EMO) reuses. It does not gain lazy-read machinery; nodetails_cached_atsemantics ship in v1.financial_metricscoverage is honest and bounded: ~63% of reports have XBRL at all, and only the 5-year window is parsed. Coverage per company is reportable (missing ≠ zero); consumers filtering on financials must treat absent metrics as "not filed digitally / outside window", never as a value.- The XBRL→nøgletal extractor is a versioned contract: Danish taxonomies change yearly, so extraction is per-taxonomy-version, and an extractor bump triggers re-parse from GCS, stamped by extractor version.
- Filing-season bursts make delta run duration elastic (thousands of documents on peak days); the run must be resumable and rate-polite rather than time-boxed.
- The delta and the metrics fetch are one run: a document-fetch failure fails the run without advancing the watermark (ADR-0005 discipline), so metrics can lag metadata only within a run, never silently across runs.
Open questions¶
- Real fetch tolerance and document sizes on
regnskaber.virk.dk(measure in the tracer bullet; G6). - G4 — updated 2026-07-19, not fully closed (issue #54,
taxonomy spike):
promoted set is
assets,equity,result,gross_profit,revenue,ebit,ebitda(candidatefsa:element names per column in the spike note); extraction approach decided as a narrow hand-rolled extractor, validated offline per taxonomy version against an Arelle-resolved sample rather than run as a runtime dependency.revenue/ebit/ebitdaare expected to be structurally low-coverage (Danish ÅRL lets small filers omit them, not a parsing bug). Residual open item: none of this is verified against a real filing yet — run the spike's ~20-filing sample before thefinancial_metricsmigration is written, per the note. - Whether
sidstOpdateretmoves on correction re-publication (G7 residual; verify during the tracer bullet — the reconciliation path covers misses either way).
Status update (2026-07-20, issue #58 — epic closure)¶
All five R0–R4 slices (#54–#58) shipped. Every genuinely new framework
capability this ADR called for landed: the no-credential adapter variant,
the load_id/ES-_id key (G5, closed), the document-fetch capability
(core/doc_fetch.py), the versioned XBRL extractor, the bounded backfill,
and reconciliation + read contract reusing (and, for _run's
pre_swap_hook, further generalizing) the CVR framework unchanged. The
financial-band read-contract view intentionally covers a single-period
band filter, not PROJECT_CHARTER.md's literal multi-year "positive
gross-profit trend" — recorded as a residual scope gap in
docs/reference/read-contract.md, not silently claimed as done.
What remains genuinely open, unresolved by this epic: every item above
this note that depends on a live filing or a live Cloud Run execution —
regnskaber.virk.dk fetch tolerance/sizes (G6), the taxonomy element-name
table's verification against ~20 real filings (G4 residual), whether
sidstOpdateret moves on correction re-publication (G7 residual), and the
first real end-to-end run of delta + eager-metrics + reconcile against
production traffic. None of these were fabricated as "measured" or
"verified" anywhere in this epic — every fixture and taxonomy claim carries
an explicit constructed/unverified flag, consistently, across #54–#58.
Amendment 2026-08-12 — Inline XBRL and whole-enterprise facts (#176, #177)¶
The extractor remains narrow and hand-rolled, but its input contract includes
Inline XBRL 2008 and 2013. Its fact QName is read from name; numeric
scale, sign, and the observed Danish ixt:numdotcomma transformation are
applied before promotion. A candidate context must have the report period end
and no segment or scenario, because either location makes it a dimensional
breakdown rather than a whole-enterprise total. Repeated facts with the same
promoted value are one observation; conflicting values remain ambiguous and
are left null. This is extractor version v2, so affected raw snapshots must
be reparsed by the bounded metrics backfill.
Amendment 2026-08-12 — the statement, not seven figures, and derived vs disclosed (#188)¶
Decision #3 promoted a small filterable set of nøgletal on the reasoning that everything else stays in the raw snapshot. Measured against a commercial product, that set was too small to be competitive: comparing our output for LASSO X A/S (CVR 34580820) against the income statement, balance sheet and nøgletal proff.dk publishes for the same company showed every figure they display is already inside the XBRL document this pipeline fetches, parses and stores. We were promoting 6 usable facts out of roughly 45. The cost of promoting the rest is a wider column list over data already in hand; the cost of not promoting them is that every consumer re-parses the raw snapshot to ask an ordinary question.
The promoted set therefore widens to the statement itself: the balance sheet
(liabilities_and_equity, contributed_capital, retained_earnings,
noncurrent_assets, current_assets, cash, trade_receivables,
provisions, shortterm_liabilities, longterm_liabilities), the income
statement (staff_costs, depreciation, profit_before_tax, tax_expense)
and employees — an exact headcount where companies.employees_band only
ever gives a band.
A column holds a fact the company disclosed. Nothing promoted is summed,
inferred, or filled from an accounting identity. Two reasons. "Missing !=
zero" only means something if a null says "the filing did not say this", and a
derived value written into the same column is indistinguishable from a
disclosed one afterwards. And a figure frozen at parse time goes stale the
moment its formula is corrected, while a view recomputes for free. So
longterm_liabilities stays null in the 78% of filings that omit it, even
though it is recoverable as liabilities_and_equity - equity - provisions -
shortterm_liabilities; a consumer that wants the identity can apply it, with
the derivation visible.
Consequently ebitda stops being a stored column. It is a pure function of
two disclosed facts and is computed by the financial_key_figures view along
with soliditetsgrad, likviditetsgrad 1, afkastningsgrad, egenkapitalens
forrentning, kapacitetsgrad, dækningsgrad and overskudsgrad. Each formula was
verified to reproduce proff.dk's published figures exactly for the filing
above. Dækningsgrad and overskudsgrad divide by turnover and are therefore
null for most Danish companies — proff.dk leaves them blank for the same
reason.
EBITDA's definition follows the market rather than the textbook (#187). The only depreciation concept Danish filings carry bundles impairment losses in with depreciation and amortisation, and proff.dk computes EBITDA from exactly that figure. Being comparable with the products a customer already uses beats being pure about impairment; the deviation is recorded here rather than hidden in a column comment.
This is extractor version v3, so raw snapshots parsed by v1 or v2 must
be reparsed by the bounded metrics backfill.
Amendment 2026-08-12 — publication metadata drops raw, and retention is bounded to the metrics window¶
This amendment changes decision 5 (publication metadata rows keep a raw
JSONB column) and decision 2 (metadata is population-wide), and follows
ADR-0005's 2026-08-12 amendment,
which removes raw from companies and production_units on measured
evidence. Decision 3's bounded five-year nøgletal window is unchanged —
this amendment aligns metadata retention to that window rather than
altering it.
Decision.
- Drop the
rawcolumn fromfinancial_reports. Decision 5's reasoning — "the metadata doc is small" — is true per document (mean 758 bytes) and false in aggregate, because this is the highest-row-count table in the hub at 6.42M publications. The publication metadata document already lands in GCS:ingesters/regnskab/pipeline.pyreusesingesters.cvr.pipeline's_run/_run_deltaunchanged, so the same run-scoped NDJSON extract is written here as for CVR entities. Decision 5's own principle — raw snapshots live in GCS, the extractor is versioned and re-runnable from them — is extended to the metadata document rather than contradicted. - Retain publication metadata only within the five-year window, keyed on
period_end >= cutoff— the same predicate and the same column_select_reports_needing_metricsalready uses, so the retained set is exactly the set eligible for nøgletal. 1,666,945 of 6,422,033 publications qualify at today's cutoff. This replaces decision 2's "all 6.41M publication records" and narrows decision 4: corrections still append andfinancial_reports_lateststill resolves them, but filing history is preserved only inside the window.
Evidence (measured, same method as the ADR-0005 amendment):
| table | rows retained | bytes/row | size |
|---|---|---|---|
financial_reports with raw, all history |
6,422,033 | 1,092 | 7.02 GB |
financial_reports without raw, all history |
6,422,033 | 170 | 1.10 GB |
financial_reports without raw, 5-year window |
1,666,945 | 170 | 0.28 GB |
financial_metrics (post-#188 columns), 5-year window |
~1.05M (63% have XBRL) | 336 | 0.35 GB |
Note the honest shape of the second lever: dropping raw is worth 5.9 GB;
bounding retention on top of it is worth a further 0.8 GB. The window was
agreed as a size measure when the hub was believed to be far larger. It is
retained deliberately, for consistency with the metrics window rather than for
the bytes.
Consequences.
- The four-entity hub lands near 2.6 GB (companies 0.49, production_units 0.62, financial_reports 0.28, financial_metrics 0.35, people + roles 0.84). Allow roughly 5 GB in service, for bloat, WAL, and autovacuum lag.
- Pre-2021 filing metadata leaves Postgres. It stays in the GCS extracts
and remains fully re-scrollable from
offentliggoerelser, which is credential-free and has no bounded history. Widening the window is still "a re-run with an earlier cutoff", exactly as decision 3 promises — but it is now a re-run of the metadata backfill as well as the metrics one. - A retained publication whose accounting period ages out of the window
becomes eligible for deletion, and its
financial_metricsrow goes with it. The FK fromfinancial_metrics.load_idhas no cascade, so the child must be deleted first or the delete fails outright — but the deeper point is that keeping derived nøgletal for a period no consumer can filter on would defeat the window. The XBRL document behind each one stays in GCS and the extractor is versioned and re-runnable against it (decision 5), so this is re-derivable, not lost. - Retention needs no scheduler of its own. The window is enforced on the
scroll — backfill, delta, and reconcile all carry it — so an out-of-window
publication is never ingested. The monthly reconcile then removes rows that
have since aged out for free: its staging+swap replaces the population with
a freshly bounded scroll, and its existing
_delete_orphaned_metricspre-swap hook takes the derived metrics with them. Steady-state retention is therefore at most one reconcile cycle stale, which is immaterial for a window measured in years. The explicitprunecommand exists for the cutover from the unbounded table and for use between reconciles. financial_reports_latestand the read contract are unaffected in shape. Consumers see fewer rows, not different columns.
Retained replay and bounded processing (#782)¶
Document acquisition returns the bytes that won the create-only write, their URI, digest, and original acquisition time. Re-extraction reads those bytes; it does not fetch a newer upstream document under an old object identity. Local and GCS document adapters follow the same rule.
Network pacing and max_fetches apply only to remote acquisition. Each run
also permits at most 10,000 snapshot reads by default. Exhaustion fails the
run after preserving completed metrics; the existing missing-extractor-version
selection resumes the next run. Candidate selection uses 250-row keyset pages
in the existing period_end desc, load_id desc order. max_docs remains an
additional caller limit. No separate progress table is introduced.