Upstream surfaces: verified facts¶
Everything on this page was probed against the live endpoints on 2026-08-11. Numbers move; the shapes and paths are the part worth trusting.
Endpoints¶
| Surface | Base URL | Search path | Auth |
|---|---|---|---|
| CVR | http://distribution.virk.dk/cvr-permanent |
<base>/<type>/_search |
Basic |
| Regnskabsdata | http://distribution.virk.dk |
<base>/offentliggoerelser/_search |
none |
The CVR base URL contains the index; the Regnskabsdata base URL does not. That difference matters, because the scroll continuation endpoint is not relative to either of them.
The scroll continuation endpoint is at the cluster root¶
POST http://distribution.virk.dk/cvr-permanent/_search/scroll -> 403 Forbidden (nginx)
POST http://distribution.virk.dk/_search/scroll -> 200
The index-prefixed path is rejected by the upstream gateway. CvrEsClient
derives /_search/scroll from the URL origin for this reason (#157) — before
that, no CVR run could read past its first 500-document page.
The delta field path differs per entity¶
CVR nests the document's last-updated field under the document's own root key. Regnskabsdata keeps it at the top level. A range query on a path that does not exist returns zero hits and no error.
| Surface / type | Documents | Delta path |
|---|---|---|
cvr-permanent/virksomhed |
2,262,072 | Vrvirksomhed.sidstOpdateret |
cvr-permanent/produktionsenhed |
2,878,638 | VrproduktionsEnhed.sidstOpdateret |
cvr-permanent/deltager |
1,824,610 | Vrdeltagerperson.sidstOpdateret |
offentliggoerelser |
6,421,900 | sidstOpdateret |
Two traps in that table:
VrproduktionsEnhedhas a capitalEin the middle.Vrproduktionsenhedmatches nothing.- 8,015
deltagerdocuments carry nosidstOpdateretinsideVrdeltagerperson. A delta can never see them; only amatch_allbackfill or reconcile picks them up.
A same-day window returns a plausible volume once the path is right:
Vrvirksomhed.sidstOpdateret >= 2026-08-10 matched 4,376 documents.
Measured throughput and memory¶
Measured from a developer machine over the public internet, loading into a
local Postgres, with raw extracts discarded. Cloud Run in europe-west1
should do no worse on the network. Documents were CVR virksomhed.
| Documents | Batch size | Wall seconds | Docs/second | Peak RSS |
|---|---|---|---|---|
| 3,000 | 500 | 6.2 | 485 | 107 MB |
| 3,000 | 1,000 | 6.6 | 452 | 155 MB |
| 3,000 | 2,000 | 6.7 | 448 | 184 MB |
| 12,000 | 1,000 | 26.8 | 448 | 157 MB |
The last two rows are the point: four times the documents at the same batch size costs the same memory. After #135, memory is a function of batch size and document size, never of the population.
Compressed extract size was ~682 gzipped bytes per virksomhed document.
At ~450 documents per second a full backfill scroll is roughly 1.4 hours for
virksomhed, 1.8 for produktionsenhed, 1.1 for deltager, and 4 hours for
offentliggoerelser — all far beyond the 600-second job timeout that was in
place when these numbers were taken (#136).
What effective-dated history cost, and what it costs now¶
The table above was taken before source history existed (#215). Adding it cost
two orders of magnitude, and #243 gave them back. All three rows below are
capped virksomhed backfills against the test project, through the transaction
pooler, on the same day:
| Run | Documents | Scroll seconds | Docs/second |
|---|---|---|---|
| Per-document history writes (#222, run 37) | 500 | 283.67 | 1.76 |
| Per-batch history writes (#243, run 38) | 500 | 5.29 | 94.47 |
| Per-batch history writes (#243, run 39) | 5,000 | 29.47 | 169.67 |
Runs 37 and 38 are the same 500 documents and wrote the identical 5,178 facts and 612 relationships, which is what makes them comparable: the semantics did not change, only the number of round trips.
The cause was never database load. The writer sent a node upsert, a fact
executemany, an edge executemany and a commit for every document, plus two
more node upserts for every edge — and one statement with its commit costs a
median 82 ms through the pooler (30 samples, select 1). The participant
loader had the same shape, at one upsert and one commit per person.
The rate still climbs with the cap because a run's fixed cost — ledger, quality verdicts, watermark — amortises over more documents. Run 39 observed 10.74 facts and 1.24 relationships per company.
243 buffers the writes per document batch. Measured by replaying fixture¶
documents through the pipeline into a local Postgres with the scroll faked, so the figures below isolate the write path and are not comparable to the end-to-end rows above:
| Documents | Round trips per document | Docs/second | Peak RSS | |
|---|---|---|---|---|
| 2,000 | per document | 6.71 | 146 | — |
| 2,000 | per batch | 0.04 | 2,141 | — |
| 12,000 | per document | 6.67 | 98 | 154 MB |
| 12,000 | per batch | 0.01 | 2,285 | 164 MB |
Both write the same 92,000 facts and 4,000 edges at 12,000 documents. Memory is still a function of batch size, not population: 164 MB against the 157 MB the same shape cost before history, because what is buffered is one batch of facts rather than a population's.
The fixture states 0.33 edges per company. Production states 1.22 (#243), so the per-document cost is nearer eleven round trips there — which is what turned a 2.26M-company scroll into a projected 356 hours.
Projection. At the measured 169.67 documents per second, and holding the
job timeouts in infra/batch_job.tf and infra/regnskab_job.tf:
| Entity | Documents | Projected | Timeout |
|---|---|---|---|
virksomhed |
2,262,072 | 3.7 h | 6 h |
produktionsenhed |
2,878,638 | 4.7 h | 6 h |
deltager |
1,824,610 | 3.0 h | 6 h |
offentliggoerelser |
6,421,900 | 4 h (unchanged) | 12 h |
Every entity fits. Three qualifications belong with that number:
- Only
virksomhedwas measured. The rate is applied to the other two CVR entities becausevirksomhedwrites the most facts per document (10.74), so they should not be slower on the write side — but neither has been measured. offentliggoerelsernever wrote source history, so #243 did not touch it and its earlier 4-hour figure stands.- The measurement is a capped run, which merges by upsert, taken from a
developer machine over the public internet. The real backfill runs in Cloud
Run
europe-west1on theCOPY+ swap path, and the rate was still climbing with cap size, so 169.67 is a conservative floor rather than a ceiling.
A capped produktionsenhed run cannot be measured before a company baseline
exists: production_units.cvr is a foreign key into companies, and a capped
scroll returns units whose owners the hub does not hold. That is the entity
ordering #140 already prescribes, not a defect — but it means the two remaining
CVR entities can only be measured during the backfill itself.
Projected hub size after the backfill¶
Everything structured lives in Supabase Postgres; only the raw source documents live in GCS. Effective-dated history is therefore a Postgres cost, and it is the largest one.
Bytes per row are measured on 2026-08-16, including indexes, and multiplied by
the populations above. The source is the clean bounded run of
cvr 10000000-10099999 — 6,152 companies, 7,349 production units, 21,243
participants and 5,927 publications — after #254 stopped bounded runs writing
the graph outside their bound and #255 slimmed the edge table:
| Table | Bytes/row | Rows | Size |
|---|---|---|---|
entity_facts |
369 | 38.1M | 14.1 GB |
entity_relationships |
291 | ~26.9M | 7.8 GB |
entity_nodes |
180 | 7.0M | 1.2 GB |
production_units |
639 | 2,879,672 | 1.8 GB |
companies |
650 | 2,262,795 | 1.5 GB |
financial_reports |
652 | 1,667,400 | 1.1 GB |
financial_metrics |
~600 | 1,667,400 | 1.0 GB |
people |
108 | 1,825,195 | 0.2 GB |
| Total | ~29 GB |
entity_relationships fell from 397 bytes a row to 291 — measured after
255, on the clean re-run of cvr 10000000-10099999. That is the 124-byte¶
saving the change was estimated at, arriving as 106. The rest of the table moved by a few bytes either way because these are smaller tables than the ones the earlier figures came from, and per-row index overhead is worst at small sizes.
How to scale a bounded measurement, and how not to¶
The first attempt at this number took the 986 MB the bounded run produced, divided by the block's 1.535% share of the register, and reported ~64 GB. That is wrong, and the way it is wrong is worth keeping.
84.7% of the edges a bounded run wrote named companies outside the bound
(#254). A participant states every company they serve, and the scroll filter
cannot reach inside the document to stop it, so a bounded run pulled 393,947
out-of-bound companies into entity_nodes as bare edge endpoints. Scaling the
whole database by register share multiplies an artifact that only exists in
bounded runs — at full population every one of those companies is in scope and
counted once.
The table above is built per node kind instead. Facts scale with nodes: 187,630 fact rows over 34,744 in-scope nodes is 5.40 per node, against 6.97M nodes at full population.
Role edges need care in the other direction. A bounded run now records only
1.24 per participant, because #254 drops the edges naming companies outside the
subset — but at full population there is no outside, so every one of those edges
is in scope and counted once. The projection therefore uses the unbounded
rate of 11.84 edges per participant, measured on 2026-08-15 before the bound was
enforced (1,040,107 edges over 87,875 participants). Using the bounded 1.24 here
would understate the table nearly tenfold, which is the same class of error as
scaling a bounded run's total by its share of the register.
operates-production-unit scales with units, at 1.11.
Two things still bound the accuracy. The 10000000-10999999 block is old
companies, and it carries about twice the participants per company of a
typical block: 2.48, against 1.57 for 25xxxxxx, 1.29 for 35xxxxxx and 1.22
for 42xxxxxx (counted live, 2026-08-16). Role history is deeper there too. So
~32 GB is an upper estimate on the edge side. Against that, financial_metrics
has never completed a backfill, so its row cost is inferred from its column
widths rather than measured.
Edges do not scale with production units. A operates-production-unit edge
is stated from both the company side (penheder) and the unit side
(virksomhedsrelation), and the two collapse on the unique version index —
measured 7,154 rows against 7,154 distinct keys. The production-unit backfill
adds none.
Most of that is overhead, not data¶
Inside entity_facts, per row:
| Bytes/row | |
|---|---|
value — the fact itself |
26 |
raw_snapshot_uri — the same GCS path repeated on every row |
73 |
rest of the heap (tuple header, source, fact_type, dates, timestamps) |
100 |
| indexes | 166 |
The stored data is 7% of the table. A repeated URI string costs nearly three times what the values do, and indexes cost more than the whole heap. Separately, 65.5% of fact rows are the only version of their fact — they restate the current projection with a validity period rather than recording a change.
entity_relationships was worse, and differently so. Per row, before #255:
| Bytes/row | |
|---|---|
the edge itself — subject, type, object, valid_from, valid_to |
36 |
attributes — a jsonb holding a fixed three-key tuple |
78 |
rest of the heap (tuple header, source, source_class, source_record_id, collected_at, raw_snapshot_uri, padding) |
92 |
| indexes (4) | 191 |
The edge was 9% of its own row. And unlike entity_facts, the table was not
versioning anything: versions per distinct edge measured 1.01 for holds-role,
1.00 for has-voting-rights-in and operates-production-unit, 1.08 for
legal-owner-of, 1.07 for beneficial-owner-of. It was an edge list paying a
bitemporal table's overhead. #255 unfolds attributes into four columns, drops
the surrogate key and its 27 MB index, and makes relationship_type and
source_class enums — about 124 bytes of the 397.
entity_facts is left alone deliberately. It does version: status 2.68,
main_industry 2.05, address 2.00, name 1.45.
These figures are measured on tables holding tens of thousands to low millions of rows, where per-row index overhead is at its worst. Treat the total as an estimate with a wide band, not a quotation.
Decision: history stays in Postgres¶
Moving entity_facts, entity_relationships, and entity_nodes out to GCS was
measured on 2026-08-15 and rejected. The numbers were good — Parquet with
zstd compresses this data 28.7× against Postgres (13.0 bytes/row against 374),
putting the whole fact history at 0.62 GB, with a per-company json.gz serving
copy at 1.28 GB and a 118 ms median lookup.
It was rejected on friction, not on cost or performance. Serving from Parquet
directly is not viable — a row group is the smallest readable unit, so every
lookup decodes ~10,000 rows to return 11, measured at 743 ms median against
118 ms for a per-company object. Splitting into an analytical shape and a
serving shape means two representations to keep in step, company_detail moves
from a PostgREST function to a service, and the unique index on
(node_id, fact_type, valid_from) — which is what makes replay converge — has
to be replaced by compaction or an open table format.
Roughly 27 GB of Supabase disk is cheaper than that complexity. Keep the history
in Postgres. If the bill ever justifies revisiting, the cheap levers are inside
this design and need no new system: raw_snapshot_uri is 73 bytes per row of
repeated string that could be a foreign key, and the indexes cost more than the
whole heap.