Skip to content

Regnskabsdata taxonomy spike — promoted financial_metrics column set

Date: 2026-07-19 Question this supports: ADR-0007's G4 — which nøgletal get promoted into financial_metrics columns, what XBRL taxonomy element(s) each maps to, and whether extraction should use a general XBRL library or a narrow hand-rolled lookup. Issue: #54 (R0 of PRD #28)

Evidence status — read this first

No real XBRL filing was fetched or parsed for this spike. This dev sandbox could not reach regnskaber.virk.dk this session (see [[regnskabsdata-ingester-plan-status]] for the network-access finding — not proven to be a blanket restriction, just unresolved this session). Everything below is built from documented knowledge of the Danish Årsregnskabsloven (ÅRL) / Erhvervsstyrelsen "Regnskab 2.0" XBRL taxonomy — the same publicly-documented taxonomy linked from #55's confirmed page fetch (https://erhvervsstyrelsen.dk/vejledning-teknisk-vejledning-og-dokumentation-regnskab-20-taksonomier-aktuelle) — not live-verified against a real instance document. Confidence is noted per-item below. This note does not close G4 outright; it narrows it to one concrete follow-up: run this mapping against ~20 real filings (spanning company sizes and a few taxonomy years) before the financial_metrics migration is written, exactly as ADR-0007's own G4 language anticipated. The natural place to do that is the first live Cloud Run run of #55/#56, or a manual one-off spike against a handful of downloaded filings if network access resolves sooner.

Background: why coverage is the central constraint, not element names

Danish accounting law (ÅRL §32) lets the smallest reporting class (Klasse B / micro) omit net turnover (revenue) from their published income statement — competitively sensitive for small firms — and start the P&L at gross profit/loss instead. Since the large majority of the ~6.4M publications in offentliggoerelser are Klasse B/micro filers (documented Danish company-size distribution: most Danish companies are small), this is not a corner case — revenue will be systematically null for most filings, not missing at random. This is why ADR-0007's candidate list already includes both revenue and gross_profit: gross_profit is the more-often-present top-line figure for the corpus as a whole. Confidence: high — this is a well-documented feature of Danish reporting-class rules, not a taxonomy-parsing detail.

The corollary for ebit/ebitda: Danish GAAP does not mandate a separate "result of primary operations" (EBIT-equivalent) line for every reporting class either — smaller filers often report straight from gross profit to net result via financial items, with no operating-result breakout tagged at all. EBITDA specifically has no single taxonomy element — it is always a derived figure (operating result + depreciation/amortisation add-back), so it inherits the operating-result coverage gap and requires a second tagged fact (depreciation/amortisation expense) to be present too. Treat ebitda as the lowest-coverage, highest-effort column in the promoted set. Confidence: medium-high — consistent with how Danish GAAP financial statement layouts are documented to work, not independently filing-verified.

Proposed promoted column set

Column Candidate taxonomy element(s) (local name, fsa: module) Coverage expectation Confidence
assets Assets (balancesum / aktiver i alt) Universal — every balance sheet has a total High
equity Equity (egenkapital i alt) Universal — required balance sheet line High
result ProfitLoss (årets resultat) Universal — the one figure every filer must disclose regardless of class High
gross_profit GrossProfitLoss (bruttofortjeneste/-tab) High — the de facto top line for Klasse B/micro filers who omit revenue Medium
revenue Revenue (nettoomsætning) Low, structurally — legally omissible for Klasse B/micro (see above); expect null on the majority of rows Medium
ebit ProfitLossFromOrdinaryOperatingActivities (resultat af primær drift) Low-medium — only filers who break out operating result separately tag this Medium
ebitda derived: ebit + DepreciationAmortisationExpense (af- og nedskrivninger) Lowest — inherits ebit's gap, needs a second fact present too Medium-low

Element local names are used above deliberately, not fully-qualified {namespace-uri}LocalName pairs: Erhvervsstyrelsen versions the taxonomy's namespace URI per release year (a new "Regnskab 2.0" version ships periodically to track ÅRL/XBRL spec changes), but the working assumption — not verified this session — is that core FSA-module concept local names stay stable release-to-release for comparability, per how XBRL taxonomy evolution is normally managed. If that assumption is wrong for any of the seven elements above, the spike-against-real-filings follow-up will surface it directly (a local name simply won't be found in an older/newer instance, forcing a per-version alias entry). Do not treat the local names in this table as final until that follow-up runs.

Extraction approach: narrow lookup, informed by a one-time Arelle pass

Decision: ship a narrow, hand-rolled extractor at runtime; use Arelle (or equivalent) only as an offline, one-time-per-taxonomy-version tool to derive and validate the local-name/context-filtering table above — not as a runtime dependency.

Reasoning:

  • The target set is small and fixed (7 columns), not full statement reconstruction — we don't need calculation-linkbase validation, dimensional roll-up resolution, or presentation-tree traversal, which is most of what a general processor like Arelle buys you.
  • The eager-on-delta design (ADR-0007) parses inline during the daily delta run, at up to a few thousand documents/day with filing-season bursts — a full XBRL processor's per-document taxonomy-loading and validation cost (which typically dominates runtime, not the fact extraction itself) is the wrong latency profile for that path. Confidence: medium — reasoned from Arelle's known architecture (loads and validates the full DTS per document), not benchmarked against this specific corpus.
  • The genuine risk a narrow extractor takes on is dimensional misattribution: an XBRL instance can carry the same concept in multiple <xbrli:context> blocks (current year vs. prior-year comparative, consolidated vs. parent-only, segment breakdowns via <xbrli:segment> dimensions). A hand-rolled extractor must filter contexts explicitly — concretely: prefer the context whose period end matches the report's own regnskab.regnskabsperiode.slutDato and whose <xbrli:segment> is absent (i.e. the whole-enterprise total, not a dimensional breakdown) — and skip the fact entirely (leave the column null) rather than guess when more than one candidate context remains. This is exactly the class of bug Arelle would catch for free via full dimensional-context resolution, which is why it stays useful as the offline validation tool: run it once per taxonomy version against the ~20-filing spike sample, confirm the narrow extractor's context-filtering picks the same values Arelle resolves, and only then trust the narrow extractor in production for that taxonomy version.
  • A taxonomy-version bump (yearly) means re-running the offline Arelle-validated spike against a fresh sample before trusting the narrow extractor for that version's filings — cheap because it's a one-time check per version, not a per-document cost.

This closes the "narrow lookup vs. general library" half of G4 with a reasoned decision; it does not close the "are these seven element names actually correct" half — that half stays open until the live-filing spike runs, as noted above.

ADR-0007 disposition

G4 updated (not fully closed) in ADR-0007: promoted column set and extraction approach are decided here; live-filing verification of the element-name table remains the residual open item, explicitly deferred to the first live run rather than guessed at further in this sandbox.