Skip to content

Upstream surfaces: verified facts

Everything on this page was probed against the live endpoints on 2026-08-11. Numbers move; the shapes and paths are the part worth trusting.

Endpoints

Surface Base URL Search path Auth
CVR http://distribution.virk.dk/cvr-permanent <base>/<type>/_search Basic
Regnskabsdata http://distribution.virk.dk <base>/offentliggoerelser/_search none

The CVR base URL contains the index; the Regnskabsdata base URL does not. That difference matters, because the scroll continuation endpoint is not relative to either of them.

The scroll continuation endpoint is at the cluster root

POST http://distribution.virk.dk/cvr-permanent/_search/scroll   -> 403 Forbidden (nginx)
POST http://distribution.virk.dk/_search/scroll                 -> 200

The index-prefixed path is rejected by the upstream gateway. CvrEsClient derives /_search/scroll from the URL origin for this reason (#157) — before that, no CVR run could read past its first 500-document page.

The delta field path differs per entity

CVR nests the document's last-updated field under the document's own root key. Regnskabsdata keeps it at the top level. A range query on a path that does not exist returns zero hits and no error.

Surface / type Documents Delta path
cvr-permanent/virksomhed 2,262,072 Vrvirksomhed.sidstOpdateret
cvr-permanent/produktionsenhed 2,878,638 VrproduktionsEnhed.sidstOpdateret
cvr-permanent/deltager 1,824,610 Vrdeltagerperson.sidstOpdateret
offentliggoerelser 6,421,900 sidstOpdateret

Two traps in that table:

  • VrproduktionsEnhed has a capital E in the middle. Vrproduktionsenhed matches nothing.
  • 8,015 deltager documents carry no sidstOpdateret inside Vrdeltagerperson. A delta can never see them; only a match_all backfill or reconcile picks them up.

A same-day window returns a plausible volume once the path is right: Vrvirksomhed.sidstOpdateret >= 2026-08-10 matched 4,376 documents.

Measured throughput and memory

Measured from a developer machine over the public internet, loading into a local Postgres, with raw extracts discarded. Cloud Run in europe-west1 should do no worse on the network. Documents were CVR virksomhed.

Documents Batch size Wall seconds Docs/second Peak RSS
3,000 500 6.2 485 107 MB
3,000 1,000 6.6 452 155 MB
3,000 2,000 6.7 448 184 MB
12,000 1,000 26.8 448 157 MB

The last two rows are the point: four times the documents at the same batch size costs the same memory. After #135, memory is a function of batch size and document size, never of the population.

Compressed extract size was ~682 gzipped bytes per virksomhed document.

At ~450 documents per second a full backfill scroll is roughly 1.4 hours for virksomhed, 1.8 for produktionsenhed, 1.1 for deltager, and 4 hours for offentliggoerelser — all far beyond the 600-second job timeout that was in place when these numbers were taken (#136).

What effective-dated history cost, and what it costs now

The table above was taken before source history existed (#215). Adding it cost two orders of magnitude, and #243 gave them back. All three rows below are capped virksomhed backfills against the test project, through the transaction pooler, on the same day:

Run Documents Scroll seconds Docs/second
Per-document history writes (#222, run 37) 500 283.67 1.76
Per-batch history writes (#243, run 38) 500 5.29 94.47
Per-batch history writes (#243, run 39) 5,000 29.47 169.67

Runs 37 and 38 are the same 500 documents and wrote the identical 5,178 facts and 612 relationships, which is what makes them comparable: the semantics did not change, only the number of round trips.

The cause was never database load. The writer sent a node upsert, a fact executemany, an edge executemany and a commit for every document, plus two more node upserts for every edge — and one statement with its commit costs a median 82 ms through the pooler (30 samples, select 1). The participant loader had the same shape, at one upsert and one commit per person.

The rate still climbs with the cap because a run's fixed cost — ledger, quality verdicts, watermark — amortises over more documents. Run 39 observed 10.74 facts and 1.24 relationships per company.

243 buffers the writes per document batch. Measured by replaying fixture

documents through the pipeline into a local Postgres with the scroll faked, so the figures below isolate the write path and are not comparable to the end-to-end rows above:

Documents Round trips per document Docs/second Peak RSS
2,000 per document 6.71 146 —
2,000 per batch 0.04 2,141 —
12,000 per document 6.67 98 154 MB
12,000 per batch 0.01 2,285 164 MB

Both write the same 92,000 facts and 4,000 edges at 12,000 documents. Memory is still a function of batch size, not population: 164 MB against the 157 MB the same shape cost before history, because what is buffered is one batch of facts rather than a population's.

The fixture states 0.33 edges per company. Production states 1.22 (#243), so the per-document cost is nearer eleven round trips there — which is what turned a 2.26M-company scroll into a projected 356 hours.

Projection. At the measured 169.67 documents per second, and holding the job timeouts in infra/batch_job.tf and infra/regnskab_job.tf:

Entity Documents Projected Timeout
virksomhed 2,262,072 3.7 h 6 h
produktionsenhed 2,878,638 4.7 h 6 h
deltager 1,824,610 3.0 h 6 h
offentliggoerelser 6,421,900 4 h (unchanged) 12 h

Every entity fits. Three qualifications belong with that number:

  • Only virksomhed was measured. The rate is applied to the other two CVR entities because virksomhed writes the most facts per document (10.74), so they should not be slower on the write side — but neither has been measured.
  • offentliggoerelser never wrote source history, so #243 did not touch it and its earlier 4-hour figure stands.
  • The measurement is a capped run, which merges by upsert, taken from a developer machine over the public internet. The real backfill runs in Cloud Run europe-west1 on the COPY + swap path, and the rate was still climbing with cap size, so 169.67 is a conservative floor rather than a ceiling.

A capped produktionsenhed run cannot be measured before a company baseline exists: production_units.cvr is a foreign key into companies, and a capped scroll returns units whose owners the hub does not hold. That is the entity ordering #140 already prescribes, not a defect — but it means the two remaining CVR entities can only be measured during the backfill itself.

Projected hub size after the backfill

Everything structured lives in Supabase Postgres; only the raw source documents live in GCS. Effective-dated history is therefore a Postgres cost, and it is the largest one.

Bytes per row are measured on 2026-08-16, including indexes, and multiplied by the populations above. The source is the clean bounded run of cvr 10000000-10099999 — 6,152 companies, 7,349 production units, 21,243 participants and 5,927 publications — after #254 stopped bounded runs writing the graph outside their bound and #255 slimmed the edge table:

Table Bytes/row Rows Size
entity_facts 369 38.1M 14.1 GB
entity_relationships 291 ~26.9M 7.8 GB
entity_nodes 180 7.0M 1.2 GB
production_units 639 2,879,672 1.8 GB
companies 650 2,262,795 1.5 GB
financial_reports 652 1,667,400 1.1 GB
financial_metrics ~600 1,667,400 1.0 GB
people 108 1,825,195 0.2 GB
Total ~29 GB

entity_relationships fell from 397 bytes a row to 291 — measured after

255, on the clean re-run of cvr 10000000-10099999. That is the 124-byte

saving the change was estimated at, arriving as 106. The rest of the table moved by a few bytes either way because these are smaller tables than the ones the earlier figures came from, and per-row index overhead is worst at small sizes.

How to scale a bounded measurement, and how not to

The first attempt at this number took the 986 MB the bounded run produced, divided by the block's 1.535% share of the register, and reported ~64 GB. That is wrong, and the way it is wrong is worth keeping.

84.7% of the edges a bounded run wrote named companies outside the bound (#254). A participant states every company they serve, and the scroll filter cannot reach inside the document to stop it, so a bounded run pulled 393,947 out-of-bound companies into entity_nodes as bare edge endpoints. Scaling the whole database by register share multiplies an artifact that only exists in bounded runs — at full population every one of those companies is in scope and counted once.

The table above is built per node kind instead. Facts scale with nodes: 187,630 fact rows over 34,744 in-scope nodes is 5.40 per node, against 6.97M nodes at full population.

Role edges need care in the other direction. A bounded run now records only 1.24 per participant, because #254 drops the edges naming companies outside the subset — but at full population there is no outside, so every one of those edges is in scope and counted once. The projection therefore uses the unbounded rate of 11.84 edges per participant, measured on 2026-08-15 before the bound was enforced (1,040,107 edges over 87,875 participants). Using the bounded 1.24 here would understate the table nearly tenfold, which is the same class of error as scaling a bounded run's total by its share of the register. operates-production-unit scales with units, at 1.11.

Two things still bound the accuracy. The 10000000-10999999 block is old companies, and it carries about twice the participants per company of a typical block: 2.48, against 1.57 for 25xxxxxx, 1.29 for 35xxxxxx and 1.22 for 42xxxxxx (counted live, 2026-08-16). Role history is deeper there too. So ~32 GB is an upper estimate on the edge side. Against that, financial_metrics has never completed a backfill, so its row cost is inferred from its column widths rather than measured.

Edges do not scale with production units. A operates-production-unit edge is stated from both the company side (penheder) and the unit side (virksomhedsrelation), and the two collapse on the unique version index — measured 7,154 rows against 7,154 distinct keys. The production-unit backfill adds none.

Most of that is overhead, not data

Inside entity_facts, per row:

Bytes/row
value — the fact itself 26
raw_snapshot_uri — the same GCS path repeated on every row 73
rest of the heap (tuple header, source, fact_type, dates, timestamps) 100
indexes 166

The stored data is 7% of the table. A repeated URI string costs nearly three times what the values do, and indexes cost more than the whole heap. Separately, 65.5% of fact rows are the only version of their fact — they restate the current projection with a validity period rather than recording a change.

entity_relationships was worse, and differently so. Per row, before #255:

Bytes/row
the edge itself — subject, type, object, valid_from, valid_to 36
attributes — a jsonb holding a fixed three-key tuple 78
rest of the heap (tuple header, source, source_class, source_record_id, collected_at, raw_snapshot_uri, padding) 92
indexes (4) 191

The edge was 9% of its own row. And unlike entity_facts, the table was not versioning anything: versions per distinct edge measured 1.01 for holds-role, 1.00 for has-voting-rights-in and operates-production-unit, 1.08 for legal-owner-of, 1.07 for beneficial-owner-of. It was an edge list paying a bitemporal table's overhead. #255 unfolds attributes into four columns, drops the surrogate key and its 27 MB index, and makes relationship_type and source_class enums — about 124 bytes of the 397.

entity_facts is left alone deliberately. It does version: status 2.68, main_industry 2.05, address 2.00, name 1.45.

These figures are measured on tables holding tens of thousands to low millions of rows, where per-row index overhead is at its worst. Treat the total as an estimate with a wide band, not a quotation.

Decision: history stays in Postgres

Moving entity_facts, entity_relationships, and entity_nodes out to GCS was measured on 2026-08-15 and rejected. The numbers were good — Parquet with zstd compresses this data 28.7× against Postgres (13.0 bytes/row against 374), putting the whole fact history at 0.62 GB, with a per-company json.gz serving copy at 1.28 GB and a 118 ms median lookup.

It was rejected on friction, not on cost or performance. Serving from Parquet directly is not viable — a row group is the smallest readable unit, so every lookup decodes ~10,000 rows to return 11, measured at 743 ms median against 118 ms for a per-company object. Splitting into an analytical shape and a serving shape means two representations to keep in step, company_detail moves from a PostgREST function to a service, and the unique index on (node_id, fact_type, valid_from) — which is what makes replay converge — has to be replaced by compaction or an open table format.

Roughly 27 GB of Supabase disk is cheaper than that complexity. Keep the history in Postgres. If the bill ever justifies revisiting, the cheap levers are inside this design and need no new system: raw_snapshot_uri is 73 bytes per row of repeated string that could be a foreign key, and the indexes cost more than the whole heap.