Regnskabsdata taxonomy spike — promoted financial_metrics column set¶
Date: 2026-07-19
Question this supports: ADR-0007's G4 — which nøgletal get promoted into
financial_metrics columns, what XBRL taxonomy element(s) each maps to, and
whether extraction should use a general XBRL library or a narrow hand-rolled
lookup.
Issue: #54 (R0 of PRD #28)
Evidence status — read this first¶
No real XBRL filing was fetched or parsed for this spike. This dev
sandbox could not reach regnskaber.virk.dk this session (see
[[regnskabsdata-ingester-plan-status]] for the network-access finding — not
proven to be a blanket restriction, just unresolved this session). Everything
below is built from documented knowledge of the Danish Årsregnskabsloven
(ÅRL) / Erhvervsstyrelsen "Regnskab 2.0" XBRL taxonomy — the same
publicly-documented taxonomy linked from #55's confirmed page fetch
(https://erhvervsstyrelsen.dk/vejledning-teknisk-vejledning-og-dokumentation-regnskab-20-taksonomier-aktuelle) —
not live-verified against a real instance document. Confidence is noted
per-item below. This note does not close G4 outright; it narrows it to one
concrete follow-up: run this mapping against ~20 real filings (spanning
company sizes and a few taxonomy years) before the financial_metrics
migration is written, exactly as ADR-0007's own G4 language anticipated.
The natural place to do that is the first live Cloud Run run of #55/#56, or a
manual one-off spike against a handful of downloaded filings if network
access resolves sooner.
Background: why coverage is the central constraint, not element names¶
Danish accounting law (ÅRL §32) lets the smallest reporting class (Klasse
B / micro) omit net turnover (revenue) from their published income
statement — competitively sensitive for small firms — and start the P&L
at gross profit/loss instead. Since the large majority of the ~6.4M
publications in offentliggoerelser are Klasse B/micro filers (documented
Danish company-size distribution: most Danish companies are small), this
is not a corner case — revenue will be systematically null for most
filings, not missing at random. This is why ADR-0007's candidate list
already includes both revenue and gross_profit: gross_profit is the
more-often-present top-line figure for the corpus as a whole. Confidence:
high — this is a well-documented feature of Danish reporting-class
rules, not a taxonomy-parsing detail.
The corollary for ebit/ebitda: Danish GAAP does not mandate a separate
"result of primary operations" (EBIT-equivalent) line for every reporting
class either — smaller filers often report straight from gross profit to
net result via financial items, with no operating-result breakout tagged
at all. EBITDA specifically has no single taxonomy element — it is
always a derived figure (operating result + depreciation/amortisation add-back),
so it inherits the operating-result coverage gap and requires a second
tagged fact (depreciation/amortisation expense) to be present too. Treat
ebitda as the lowest-coverage, highest-effort column in the promoted set.
Confidence: medium-high — consistent with how Danish GAAP financial
statement layouts are documented to work, not independently filing-verified.
Proposed promoted column set¶
| Column | Candidate taxonomy element(s) (local name, fsa: module) |
Coverage expectation | Confidence |
|---|---|---|---|
assets |
Assets (balancesum / aktiver i alt) |
Universal — every balance sheet has a total | High |
equity |
Equity (egenkapital i alt) |
Universal — required balance sheet line | High |
result |
ProfitLoss (årets resultat) |
Universal — the one figure every filer must disclose regardless of class | High |
gross_profit |
GrossProfitLoss (bruttofortjeneste/-tab) |
High — the de facto top line for Klasse B/micro filers who omit revenue | Medium |
revenue |
Revenue (nettoomsætning) |
Low, structurally — legally omissible for Klasse B/micro (see above); expect null on the majority of rows | Medium |
ebit |
ProfitLossFromOrdinaryOperatingActivities (resultat af primær drift) |
Low-medium — only filers who break out operating result separately tag this | Medium |
ebitda |
derived: ebit + DepreciationAmortisationExpense (af- og nedskrivninger) |
Lowest — inherits ebit's gap, needs a second fact present too |
Medium-low |
Element local names are used above deliberately, not fully-qualified
{namespace-uri}LocalName pairs: Erhvervsstyrelsen versions the taxonomy's
namespace URI per release year (a new "Regnskab 2.0" version ships
periodically to track ÅRL/XBRL spec changes), but the working assumption —
not verified this session — is that core FSA-module concept local
names stay stable release-to-release for comparability, per how XBRL
taxonomy evolution is normally managed. If that assumption is wrong for any
of the seven elements above, the spike-against-real-filings follow-up will
surface it directly (a local name simply won't be found in an older/newer
instance, forcing a per-version alias entry). Do not treat the local names
in this table as final until that follow-up runs.
Extraction approach: narrow lookup, informed by a one-time Arelle pass¶
Decision: ship a narrow, hand-rolled extractor at runtime; use Arelle (or equivalent) only as an offline, one-time-per-taxonomy-version tool to derive and validate the local-name/context-filtering table above — not as a runtime dependency.
Reasoning:
- The target set is small and fixed (7 columns), not full statement reconstruction — we don't need calculation-linkbase validation, dimensional roll-up resolution, or presentation-tree traversal, which is most of what a general processor like Arelle buys you.
- The eager-on-delta design (ADR-0007) parses inline during the daily delta run, at up to a few thousand documents/day with filing-season bursts — a full XBRL processor's per-document taxonomy-loading and validation cost (which typically dominates runtime, not the fact extraction itself) is the wrong latency profile for that path. Confidence: medium — reasoned from Arelle's known architecture (loads and validates the full DTS per document), not benchmarked against this specific corpus.
- The genuine risk a narrow extractor takes on is dimensional
misattribution: an XBRL instance can carry the same concept in multiple
<xbrli:context>blocks (current year vs. prior-year comparative, consolidated vs. parent-only, segment breakdowns via<xbrli:segment>dimensions). A hand-rolled extractor must filter contexts explicitly — concretely: prefer the context whose period end matches the report's ownregnskab.regnskabsperiode.slutDatoand whose<xbrli:segment>is absent (i.e. the whole-enterprise total, not a dimensional breakdown) — and skip the fact entirely (leave the column null) rather than guess when more than one candidate context remains. This is exactly the class of bug Arelle would catch for free via full dimensional-context resolution, which is why it stays useful as the offline validation tool: run it once per taxonomy version against the ~20-filing spike sample, confirm the narrow extractor's context-filtering picks the same values Arelle resolves, and only then trust the narrow extractor in production for that taxonomy version. - A taxonomy-version bump (yearly) means re-running the offline Arelle-validated spike against a fresh sample before trusting the narrow extractor for that version's filings — cheap because it's a one-time check per version, not a per-document cost.
This closes the "narrow lookup vs. general library" half of G4 with a reasoned decision; it does not close the "are these seven element names actually correct" half — that half stays open until the live-filing spike runs, as noted above.
ADR-0007 disposition¶
G4 updated (not fully closed) in ADR-0007: promoted column set and extraction approach are decided here; live-filing verification of the element-name table remains the residual open item, explicitly deferred to the first live run rather than guessed at further in this sandbox.