Read contracts¶
This page separates the implemented read views from the accepted company-detail
contract. Consumers, such as data-analysis, use these published surfaces.
They never read base tables or call upstream sources.
Schema reference is the one index of every surface
granted to authenticated, mapping each to its migration and to its section
here. Known consumers records which repository depends on which
surface. A consumer scopes a data feature by enumerating from those two pages
and the migrations, not from the fields its application already reads
(ADR-0019).
Company-detail contract¶
Status: implemented (#215, #216, #217, #218, #219, #220, #221, #507). The company-detail operation is the single company-selection surface.
The operation is the PostgREST function company_detail(p_cvr bigint). It is
security definer with a pinned search path over base tables that keep
row-level security with no policies, and stable, so it reads and never
writes. company_facts_as_of(p_cvr bigint, p_as_of date) answers what the
source stated on a given date.
Selecting one company by CVR returns one object with three required parts:
current_detail contains the current company, production units,
participants and roles, registry contact points, web signals, and financial
publications and metrics. Financial publications stay limited to the latest
five accounting years.
current_detail section |
Required content |
|---|---|
| company | identity, names, lifecycle and status, legal form, primary and secondary industries, locations, advertising protection, registry contacts, employment, and current company attributes |
| production units | identity, current parent, names, lifecycle and status, industries, locations, registry contacts, and employment |
| participants | participant identity and kind, current allowed name, current roles, role functions, ownership, and voting rights |
| financials | five-year cutoff, publications, correction state, document references, metrics, and absence reasons |
| web | current presence, web contacts, company-profile signals, entity mentions, and website people, each with evidence |
| freshness | source-slice state, source update time, collection time, and refresh state |
source_history contains effective-dated source facts for the company and its
related CVR entities. It is queryable without reading GCS. Each fact identifies
the entity, source, fact type and value, effective period, source update time,
collection time, and raw snapshot.
relationship_history contains current and ended edges that touch the company.
Each edge identifies both entities, relationship type, effective period, source
and trust class, observation time, and evidence. Registry edges and observed
web edges remain distinguishable.
other_entity identifies the counterpart with kind (company,
production-unit, or participant), source_key (CVR number, P-number, or
enhedsNummer), and name: the counterpart's current registry name (company
name, production-unit name, or participant name within ADR-0006). The name is
the current name, not the name at the time of the edge, so an ended role or
production unit is named too. It is null when the registry has no name.
participant_kind is the participant's person, company, or other kind
and is null for a company or production-unit counterpart (#870).
The response includes source-slice freshness and an explicit absence reason when data is missing. A missing value does not mean that the source published zero, false, or an empty relationship.
The public surfaces are:
| Surface | SQL name | Status |
|---|---|---|
| company-detail operation | company_detail(bigint) |
implemented, company slice only |
| company-search operation | search_companies(text, jsonb, text, text, jsonb) |
implemented, fixed 100-row pages with Facet filters |
| company-Facet operation | company_facet_counts(text, text, jsonb) |
implemented, exact disjunctive counts |
| company-search filter contract | company_search_filter_contract() |
implemented, self-describing accepted filter keys |
| bulk participants view | company_participants |
implemented, current registry roles only |
| graph-neighbor operation | company_graph_neighbors(bigint[], numeric, jsonb) |
implemented, current observed competitor-of and customer-of edges |
| hybrid-similarity operation | hybrid_similar_companies(bigint[], jsonb, jsonb) |
implemented, RRF over semantic and structured ranks at request time |
| as-of source read | company_facts_as_of(bigint, date), production_unit_facts_as_of(bigint, date), participant_facts_as_of(bigint, date) |
implemented |
| current-company view | active_companies |
implemented |
| current web-signal views | company_web_presence, company_web_contact, company_profile_signals, company_entity_mentions |
implemented |
| current financial views | financial_reports_latest, active_companies_financial_snapshot, financial_key_figures |
implemented |
entity_nodes, entity_facts, and entity_relationships are not consumer
surfaces. They are read through the operation, which is the only granted path
to them.
Company-search contract¶
Status: implemented (#550, spec #549).
The PostgREST operation is:
search_companies(
p_query text default null,
p_filters jsonb default '{}'::jsonb,
p_sort_by text default 'relevance',
p_sort_direction text default 'asc',
p_cursor jsonb default null
) returns jsonb
Only authenticated can execute it. The function is stable,
security definer, and has the fixed search path public, pg_temp. Base
tables stay behind row-level security.
The result has items, next_cursor, and data_version. items has at
most 100 rows. Each row contains cvr, name, domain,
main_industry, main_industry_text, municipality_code,
municipality_name, status_group, postal_code, employees_band,
started_at, and the latest financial_period_end, gross_profit,
revenue, result, equity, ebit, derived ebitda, and assets.
municipality_name is the current registry municipality name for
municipality_code, the same value as active_companies.municipality_name.
It is null when the registry has no name. status_group is the token that the
status_groups filter and company_facet_counts().facets.status_groups use
for the company: active, bankrupt, voluntary_liquidation,
compulsory_dissolution, reassumption, or unknown when the registry
status is absent or not mapped. Both fields are present in filtered and
unfiltered searches (#904); the authenticated regression
test_company_search_contract.py::test_company_search_rows_return_municipality_name_and_status_group
covers both paths and missing values.
main_industry_text is the source Danish industry label, including DB25 labels.
It is present in both filtered and unfiltered searches and is null when the
source has no label. Do not derive it from a bundled DB07 code table. The same
field is available on active_companies. The authenticated regression
test_company_search_contract.py::test_company_list_surfaces_return_source_industry_text
covers both search paths, exact-CVR search, source labels, and missing labels.
The operation starts from every active company. For the existing registry sorts, it left-joins optional financial output after selecting the candidate page. Financial ordering reads the latest parsed figures before selecting the page. A company with no resolved website or no parsed financial report stays in the result with null values. It never receives zero as a substitute.
Search order is:
- exact CVR;
- exact case-insensitive legal name or normalized website domain;
- legal-name or domain prefix;
pg_trgmfuzzy legal-name or domain match at the tested 0.3 threshold.
Without an effective filter, the existing sorts are relevance, cvr, name,
main_industry, municipality_code, postal_code, employees_band, and
started_at, in either direction. With an effective filter, the existing
registry order is name ascending.
Both filtered and unfiltered searches also accept gross_profit, revenue,
result, and employees with p_sort_direction = 'desc' (#859).
The financial sorts use the same latest parsed accounting period as the
returned financial figures: latest period end, then latest publication,
then load ID. employees uses the latest numeric CVR employment count,
not the text employment band or the average headcount from a filing.
The selected value is primary and CVR ascending is the final stable tie-breaker. Missing values sort last, including when the next page crosses into missing values. These are server-side keyset orders over all matching companies, not a sort of the first result page. An unsupported sort or direction is an invalid parameter error.
The filter object has industries, employee_groups, municipality_codes,
and status_groups. Values in one Facet use OR. Different Facets use AND.
Missing values use unknown; they never use zero. The cursor is a JSON object
that the consumer treats as opaque. It binds the normalized query, normalized
filter object, sort, last sort value, final CVR, and data_version.
data_version combines the latest successful CVR company run, the latest
successful Regnskabsdata publication run, and the latest terminal
web-enrichment run, the financial mutation revision, the current report
relation identity, and the UTC retention date. If one changes, the operation rejects the old cursor with
company search cursor expired because data changed. The consumer starts at
page one and does not mix daily versions.
tests/test_company_search_numeric_sorting.py exercises the authenticated
numeric-sort contract: latest figures, registry headcount, empty filters,
ties across pages, the transition to null values, and rejection of changed
query, filter, sort, data version, or malformed cursor values.
Company-search performance¶
The first implementation materialized all active companies before fuzzy candidate selection. It measured 35.686 seconds on 2,260,000 active synthetic companies and was rejected.
The original registry-sort query uses separate indexed paths for exact CVR, exact name, prefix, and trigram candidates. It selects at most 101 CVRs before joining financial output. These measurements do not cover numeric financial ordering, which must inspect each matching company's latest figure. On the same full population, the original query measured:
| Request | Execution time |
|---|---|
| one-letter fuzzy company-name search | 204 ms |
| default first page, sorted by name | 41 ms |
Both meet the 500 ms requirement. The reproducible opt-in test is
tests/test_company_search_performance.py with RUN_PERFORMANCE_TESTS=1.
The measurements use local synthetic data and are not a production capacity
forecast.
Company-Facet contract¶
Status: implemented (#558, parent #428). The PostgREST operation is:
company_facet_counts(
p_data_version text,
p_query text default null,
p_filters jsonb default '{}'::jsonb
) returns jsonb
Only authenticated can execute it. The result contains data_version,
calculated_at, total, and exact Facet values for industry, employee group,
municipality, status, company_form_codes, advertising_protected,
has_website, has_email, and has_phone. Search applies before every count. Each Facet ignores
its own selection and keeps the selections from all other Facets. A selected
value with zero matches stays in the result. Boolean facets always return JSON
true and false buckets, including zero counts. Other zero values are absent.
Legal-form values are source codes with source labels (or the code if the label
is absent); a missing form has no bucket. Website, email, and phone refer to
the retained web-presence/contact rows, as their filters do. Multiple source
rows for one CVR count once. Registry contact fields remain separate. A
changed data version is an invalid parameter error.
total (#617) is the exact size of the whole matched set for the request —
the number of companies search_companies would return for the same query and
p_filters, with every filter applied (all Facet selections AND-combined,
plus the registry / financial / web-presence Filterable fields). It is the
count for range and trigger filters without per-value buckets: the
consumer's "X virksomheder i målgruppen". With no status selection,
sum(status_groups[].count) equals total under the same snapshot.
Industry values are exact six-digit leaf codes and source Danish text.
Employee values are zero, 1_9, 10_49, 50_199, 200_plus, and
unknown. Municipality values use the source code and the hub's official
kommunekode to municipality-name reference; the source kommuneNavn is only
used when a code is not in that reference. Status
values are active, bankrupt, voluntary_liquidation,
compulsory_dissolution, reassumption, and unknown. The active grid does
not include ended companies.
All twelve facet groups use the complete company_facet_cube. The index stores
exact counts for shared combinations and effective-date intervals. A CVR
membership relation connects text-search and filter-only matches to those same
combinations. Each request selects the current UTC interval, so holding changes,
control changes, and financial expiry do not need a midnight publication job.
The count query groups values before building JSON and uses at most 32 MB of
working memory per query-plan operation.
Source statements update the affected memberships and counts in their transaction. Large source statements rebuild the index; an atomic company-table replacement publishes the control projection and index before commit. The CVR loader writes pages as set-based upserts and keeps the final record for a repeated CVR. An unchanged control publication does not cause a graph rebuild for each metadata update. Web-presence and contact writes also update the index and publication revision. A caller must repeat search if its data version is no longer current.
The opt-in proof uses 2.26 million synthetic companies and 20 measured samples after one warm-up per context. It includes source company forms and holding intervals, 6,000 companies with control evidence, 12,000 companies with four annual filings each, and web presence and contacts. Its cases cover broad, selective, fuzzy, presence, structure, financial-class, combined, and restored-form requests. It checks literal expected class counts and records row-at-a-time upserts, page upserts, and full replacement throughput. This controlled fixture is not a measured production distribution.
With the previous implementation, the original fixture reproduced an unfiltered p95 of 8,459.6 ms, four-facet p95 of 144.2 ms, and fuzzy-search p95 of 106.1 ms. The complete index passed all eight contexts on local PostgreSQL 17 with the populated fixture. These p95 values use 20 samples after one warm-up:
| Context | p95 |
|---|---|
| Unfiltered | 404.9 ms |
| Original four facets | 426.3 ms |
| Fuzzy search and facets | 122.4 ms |
| Website, email, and phone presence | 434.9 ms |
| Holding and group roles | 348.8 ms |
| Financial classes | 331.9 ms |
| Combined facets | 343.4 ms |
| Restored absent company form | 328.7 ms |
Source-write measurements include projection and index publication. Unchanged single-row upserts reached 370 rows/s; unchanged 1,000-row pages reached 2,247 rows/s. Pages that changed advertising-protection values reached 237 rows/s. A full replacement of 2,260,000 companies took 395.1 seconds (5,720 rows/s) and retained the exact public total. An explicit complete-index refresh took 355.6 seconds. Initial population and index creation took 214.6 seconds; control, financial, and web evidence publication took 103.7, 110.9, and 11.2 seconds. These are local measurements, not production throughput guarantees.
Run the proof on a local disposable database with all migrations applied:
DATABASE_URL=postgresql://localhost/disposable RUN_PERFORMANCE_TESTS=1 \
uv run pytest tests/test_company_facet_performance.py -q -s
Self-describing surfaces¶
Status: implemented (#629, spec #626, ADR-0019). A published surface reports its own accepted shape from the same place it enforces it, so a consumer reads the contract without reading a base table and the reported shape cannot drift from the code.
company_search_filter_contract() is the filter contract for Company search.
It is security definer, stable, fixed search path, granted to
authenticated, revoked from public. It returns one row per accepted
p_filters key:
| Column | Meaning |
|---|---|
filter_key |
The p_filters key, e.g. industries. |
value_domain |
The scalar type of each value: text, integer, numeric, date, or boolean. |
filter_shape |
set — the value is a JSON array, matched any-of; range — the value is {"min": x, "max": y} (either bound optional, at least one required), the company value must be non-null and inside; boolean — the value is JSON true / false, matched against a presence flag. |
vocabulary_kind |
set keys with a vocabulary: enum — a fixed list; reference — values from a named reference. Null when no vocabulary applies. |
enum_values |
The fixed list, when vocabulary_kind is enum; null otherwise. |
reference_name |
The reference the values come from (cvr_branchekode, kommunekode, company_form_vocabulary, dk_postnummer, cvr_ansatte_kode), when vocabulary_kind is reference; null otherwise. |
accepts_unknown_sentinel |
Whether the key also accepts the string unknown for a company whose source value is missing. |
facet_capable |
Whether company_facet_counts returns an exact disjunctive bucket list for this key. Range and trigger keys are filter-only. |
read_contract_anchor |
The section of this page a consumer reads for the key's semantics. |
The original four facet_capable keys — industries, employee_groups,
municipality_codes, status_groups — are anchored to
Company-Facet contract, are set-shaped, and all
have accepts_unknown_sentinel true: a company whose source value is missing
matches unknown, never zero. accepts_unknown_sentinel is stated per key so
a consumer validating from this contract alone knows unknown is allowed even
for a reference or integer key, where it is not in enum_values. The wider
registry set is at Filterable fields: registry.
company_form_codes, advertising_protected, has_website, has_email, and
has_phone also report facet_capable = true; their accepted filter shapes
and source meanings are unchanged.
The function assembles its answer from company_search_filterable_field, an
immutable reference table. Adding a Filterable field is one row there; the
contract and the enforcement below both follow from it.
Unknown-key and malformed-value rule. search_companies and
company_facet_counts read the same reference table. A p_filters key that is
not in it is an invalid-parameter error (22023) whose message names the key —
a consumer typo or a stale key is an error, not a silent no-op. A p_filters
value that is not a JSON object, a set value that is not an array, a
boolean value that is not true / false, or a range value that is not a
{min, max} object with at least one bound, is the same class of error. The
key set the contract reports is exactly the key set the two operations
enforce.
Company CVR set filter¶
cvrs restricts search_companies and company_facet_counts to an explicit
set of active companies (#860). A consumer can save this filter with the other
Saved view filters; the producer does not store a list.
The contract reports filter_shape = set, value_domain = integer,
accepts_unknown_sentinel = false, and this section's anchor. Supply 1–5,000
eight-digit JSON integer CVRs. Strings, fractions, nulls, booleans, values outside
10000000–99999999, an empty array, and more than 5,000 entries raise SQLSTATE
22023; the message names cvrs. Duplicate CVRs do not duplicate companies.
An inactive or unknown CVR contributes no row and does not cause an error.
The set uses OR within cvrs and AND with the query and every other filter.
Exact-CVR p_query remains an intersection, not an override. The normal
filtered name and numeric sort rules, 100-row pages, and bound cursors apply.
CVR membership narrows candidates before the other predicates. Facet totals
and disjunctive buckets remain exact within that set; there is no CVR bucket.
tests/test_company_search_filter_contract.py covers the public contract,
input bounds, intersections, paging, numeric ordering, and facet counts.
Filterable fields: registry¶
Status: implemented (#630, spec #626). Company search can filter on every
registry attribute the collection sources. Each key is a row in
company_search_filterable_field, enforced by search_companies, published by
company_search_filter_contract(), and — for a range or boolean key —
applied by company_facet_counts as a filter-only constraint on the counted
population (the pre-aggregated cube is left for a company scan when such a key
is present).
| Key | Shape | Company attribute |
|---|---|---|
company_form_codes |
set (company_form_vocabulary) |
company_form_code |
postal_codes |
set (dk_postnummer) |
postal_code |
postal_code |
range (integer) | postal_code; composes with postal_codes using AND |
secondary_industries |
set (cvr_branchekode) |
secondary_industries — array overlap, any-of |
latest_employment_bands |
set (cvr_ansatte_kode) |
latest_employment_band |
registered_capital |
range (numeric) | registered_capital (DKK) |
incorporation_date |
range (date) | started_at — company age / incorporation date |
latest_employment_count |
range (numeric) | latest_employment_count — the exact latest headcount |
advertising_protected |
boolean | advertising_protected |
has_registry_phone |
boolean | a registry phone is present |
has_registry_fax |
boolean | a registry fax is present |
has_registry_email |
boolean | a registry email is present |
has_registry_website |
boolean | a registry website is present |
missing != zero. A range filter on a nullable attribute
(postal_code, registered_capital, started_at, latest_employment_count) excludes a
company that has no value — it is never read as zero or as the epoch. A set
filter likewise matches only a company whose attribute is one of the listed
values; a null attribute is not a match. advertising_protected and the
has_registry_* flags are always determinate.
Out of scope. Profile-signal semantic fields (line_of_business,
products_or_services, value_proposition, and the rest of
web signals) are not Filterable fields — they are model
observations, not registry facts, and a filter would present a confidence-
ranked reading as a hard predicate. Participant and graph-neighbour attributes
(a person's role, an ownership percentage, a competitor edge) are answered by
company_participants and company_graph_neighbors, not by a Company-search
filter.
Filterable fields: financial¶
Status: implemented (#630, spec #626). Company search can filter on the
financial statement the collection parses. Every figure and ratio is read from
the company's single most recent parsed accounting period,
correction-resolved — the LATERAL "latest per company" shape #58 measured and
kept (see Performance (financial-band filter, #58)).
Search applies company_search_matches_cvr_keyed_filters(cvr, filters) to
candidates. Figure and ratio keys are filter-only. Facet requests first select
matching CVRs, then use their index memberships to count the facet values.
Requests without text resolve the latest financial rows as one set.
| Key | Shape | Source |
|---|---|---|
gross_profit |
range (numeric) | financial_metrics.gross_profit (bruttofortjeneste) |
revenue |
range (numeric) | financial_metrics.revenue — disclosed by a minority of filings (ADR-0007) |
equity |
range (numeric) | financial_metrics.equity |
result |
range (numeric) | financial_metrics.result (årets resultat) |
ebit |
range (numeric) | financial_metrics.ebit |
ebitda |
range (numeric) | derived: ebit + depreciation, never stored |
assets |
range (numeric) | financial_metrics.assets |
filing_employee_count |
range (numeric) | financial_metrics.employees — the exact headcount the filing reported, not latest_employment_* |
equity_ratio |
range (numeric) | soliditetsgrad: equity / assets |
current_ratio |
range (numeric) | likviditetsgrad 1: current_assets / shortterm_liabilities |
return_on_assets |
range (numeric) | afkastningsgrad: result / assets |
return_on_equity |
range (numeric) | egenkapitalens forrentning: result / equity |
gross_margin |
range (numeric) | dækningsgrad: gross_profit / revenue |
operating_margin |
range (numeric) | overskudsgrad: ebit / revenue |
has_parsed_financials |
boolean | the company has at least one XBRL-bearing, parsed report |
latest_filing_period |
range (date) | period_end of the most recent report |
The six ratios use the exact financial_key_figures formulas, each nullif()-
guarded on its denominator. gross_margin and operating_margin are null for
most companies because turnover is not disclosed.
missing != zero. A range filter needs a non-null value inside the bound.
A company with no parsed report is excluded by any figure, ratio, or
latest_filing_period filter — never read as zero. A company whose filing does
not disclose an input is excluded by a filter on that figure or on a ratio
derived from it (a null depreciation makes ebitda null; a null revenue
makes both margins null). has_parsed_financials is the one determinate
financial key.
Out of scope. Multi-year trend predicates ("positive gross-profit trend
over three years") are not Filterable fields — the surface exposes the latest
period only, the same scope note #58 recorded. A filing older than the retained
five-year window (ADR-0007) has no row at all, so latest_filing_period
filters within that window.
Web classification filters¶
The public contract accepts these source-qualified website observations:
| Key | Shape | Meaning |
|---|---|---|
sales_areas |
Set: local, regional, national, international |
Quoted service or delivery coverage with named geography. |
target_markets |
Set: b2b, b2c, b2g |
Quoted company statements about business, consumer, or public-sector customers. |
technologies |
Set: wordpress, woocommerce, shopify, google_tag_manager, google_tag, meta_pixel |
Approved installation signatures in retained HTML. |
has_webshop |
Boolean | An identified Shopify or WooCommerce product-add form. |
has_google_tag |
Boolean | A Google loader and a matching configured tag identifier. |
has_facebook_pixel |
Boolean | A Meta loader and a literal Pixel initialization call. |
ADR-0025 fixes the source rules and detector catalog. Each set uses OR within the selection. Different keys use AND. Unknown enum values are rejected. Empty sets add no condition. A false boolean means no approved signature was found in the recorded bounded HTML coverage. It does not mean absence across the entire website. Missing, failed, conflicting, expired, and identity-scoped HTML evidence cannot match either boolean selection. Tag Manager alone does not establish a Google tag.
company_detail(cvr).current_detail.web_classifications exposes each key's
current evaluation, latest_attempt, and append-only history. Evaluations
include eligibility, freshness, value, source URLs and locators, coverage,
confidence where applicable, rule/extractor/prompt/model versions, and source
fetch and expiry times. The underlying evidence table is not granted to clients.
The frozen page_fresh_seconds policy currently gives evidence 30 days from
the earliest source fetch. Re-extraction does not renew this period. A failed
refresh preserves earlier eligible evidence and records the failed attempt.
A newer successful negative or conflict supersedes earlier positives. Expired
evidence remains visible in detail but is excluded from search and exact
counts. Publication writes and expiry boundaries invalidate search versions.
Search itself does not fetch pages. Existing free-text profile values are not
converted into filter values; normal refresh or compatible re-extraction must
produce the new versioned outputs.
The consumer implementation is sales.sourceagent#187. On a retained fixture with 2,260,001 companies and 60,006 synthetic website evaluations, industry/municipality queries with Google tag, national coverage, or combined technology/tag selections took 0.021–0.049 seconds for search and 0.475–0.533 seconds for exact facets.
Company contact availability¶
The Kontakt group uses the existing company flags has_registry_phone,
has_registry_email, has_registry_website, and advertising_protected, plus
these fields:
| Key | Shape | Meaning |
|---|---|---|
registry_roles |
Set: board, executive_management |
Any current qualifying natural-person CVR membership. |
has_attributed_work_contact |
Boolean | At least one retained, unsuppressed contact attributed to a company-scoped person mention. |
has_board_work_contact |
Boolean | Such a contact with an unreversed identity link to this company's current board member. |
The role mapping and privacy boundary are fixed in
ADR-0024.
Role dates come from each member's FUNKTION, not its organisation container.
Both endpoints are inclusive on the UTC read date. Unknown participant kinds,
legal entities, and board alternates do not match. Unknown role values are
rejected. An empty set adds no condition. Other filter keys combine with AND.
company_detail(cvr).current_detail.contact_availability publishes the same
roles and nullable flags, the rule version, registry provenance, and the people
evaluation in the latest relevant run with its input coverage and versions. It
returns no personal contact values or suppression reason.
True requires retained supporting evidence. False requires a successful latest people evaluation over a nonempty recorded bounded text input and no retained supporting value. Board false also requires an evaluable registry collection. False refers to this retained dataset, not all people or pages in the world. Uncollected, failed, or unevaluable routes yield null without positive evidence; both boolean selections exclude null. Later collection failure does not remove retained positive contact evidence. Contacts have no age-based expiry under ADR-0015. Availability does not mean verified identity, verified delivery, or permission to contact someone.
Suppression, contact deletion, identity-link correction, collection results, and membership changes invalidate search versions in the same transaction. Counts remain exact and never duplicate a company for multiple contacts or roles. Normal CVR collection fills the new membership projection; older rows remain unavailable until refreshed. Search does not acquire new evidence. The existing general role-history date defect is tracked in issue #720.
Employee and decision-maker contact filters are unsupported. A website title cannot establish those relationships. The consumer must use these public keys and must not join personal-data base tables.
The consumer change is sales.sourceagent#185. Its live contract conformance suite must pass before delivery. On the retained 2,260,001-company fixture, an industry and municipality query with attributed contact presence took 0.058 seconds for search and 0.725 seconds for exact facets. The equivalent board-role query took 0.030 and 0.521 seconds. The internal presence predicate reads only the required evidence and reuses its CVR-keyed plans; building full detail provenance per candidate took 5.50 seconds for contact facets. Broader unselective query work remains in #634.
Current registry auditors¶
auditor_ids is an any-of set of strings from
company_auditor_vocabulary(). Each ID is a CVR enhedsNummer for a company
participant, not a CVR number or an audit-firm name. The vocabulary supplies
the latest disclosed name. An empty set adds no condition; a nonempty set
requires a matching current registry firm. Unknown or ended IDs return no
matches. Malformed identifiers are rejected. All other filters combine with
AND. Counts remain exact and count a company once, including joint auditors.
company_detail(cvr).current_detail.auditors contains current firms, the full
dated collection from the current document (history), and retained source
publications. Each publication includes source update, collection time, and
raw reference. state=unavailable means the relationship collection was not
collected; state=collected includes an empty collection. Empty does not claim
that no real-world auditor exists. unavailable entries explain unsupported
participant kinds, missing identity, or missing membership starts. Natural
persons remain source evidence and cannot match the firm filter.
Current means FUNKTION=REVISION membership started on or before the UTC read
date and has no end or an end on or after that date. Organisation-container
dates do not substitute for membership dates. CVR registry appointments are
separate from filing audit firms, signing persons, and submitting enterprises.
See ADR-0023.
CVR delta writes and population swaps publish appointments atomically and keep old source snapshots in the restricted auditor-publication table. Source changes and UTC date changes expire search cursors. Normal CVR refresh or reconciliation fills old rows; until then their auditor evidence is unavailable. Search never contacts CVR. No additional source or credentials are required. The sales consumer change is sales.sourceagent#183. Its live filter conformance suite must pass before release.
The opt-in auditor performance fixture contains 2,260,000 synthetic companies and auditor evidence for one fifth of them. The verified industry, municipality, and auditor query returned one company: search took 0.074 seconds and exact facets took 0.515 seconds. Before the query changes, search took 1.134 seconds and scanned 2,034,529 nonmatching active names. Binding the four facet arrays before planning lets PostgreSQL use the industry index. Reusing the internal CVR-keyed plan reduced repeated planning in exact facets. Auditor-only queries remain subject to the broader query work tracked in #634.
Gross-profit growth¶
gross_profit_growth is a growth object, for example
{"years": 3, "min": 20}. Both fields are required. years is 1 or 3;
min is a finite percentage with an inclusive lower bound. No other child
keys are accepted. company_search_filter_contract().object_schema declares
this shape and the allowed horizons. Other filter shapes have a null
object_schema.
The percentage is 100 * (latest - baseline) / baseline. A baseline must be
positive: 100 to 120 is 20%; 100 to -20 is -120%. A missing, zero, or negative
baseline is unavailable. This is endpoint change, not CAGR or a claim that
profit grew in each intervening year.
Select the latest retained accounting period, then its latest publication, before looking for metrics. A newer unparsed or PDF-only correction makes that endpoint unavailable. The baseline ends exactly N calendar years earlier; Postgres normalizes a leap day to the last day of February. Both periods must be annual and their start dates must also differ by exactly N calendar years. Both endpoints must fall within the existing five-year retention window. Missing intervening periods do not prevent comparison.
The extractor retains the fact's ISO4217 currency, entity identifier, start date, and non-dimensional scope. Both endpoints must use the same currency and entity and match their filing periods. Unverified or conflicting metadata is unavailable. No currency conversion or consolidation inference occurs. Old extraction rows become eligible only after the existing metrics backfill re-extracts them with version v5. No unit or scope is filled from a default.
The predicate applies to search_companies and every company_facet_counts
bucket and exact total. It has no separate bucket. Other keys combine with
AND; existing disjunctive facet rules remain unchanged.
company_detail(cvr).current_detail.financials.gross_profit_growth contains
entries 1 and 3. Each reports percentage, absence_reason, rule_version,
and both endpoint objects with publication ID, dates, value, comparison
metadata, source timestamps, parsing time, taxonomy/extractor version, and raw
source reference. A missing endpoint is null. Reasons distinguish missing
endpoints, unavailable metrics, nonpositive baseline, incompatible periods,
currency, scope, fact periods, nonfinite metrics, and the retention cutoff.
Financial mutations advance a derived publication revision in the same transaction. The search version also includes the UTC date to bind the rolling retention cutoff. Previous pages and facet contexts expire after either changes. Bulk replacement changes the report relation identity atomically and preserves the publication triggers for later mutations.
Migration: 20260910120000_gross_profit_growth.sql.
Decision: gross-profit endpoint comparisons.
Filterable fields: web presence¶
Status: implemented (#630, spec #626). Company search can filter on the
web-enrichment presence output. All five keys are determinate booleans applied
by company_search_matches_cvr_keyed_filters(cvr, filters). has_website also
has an exact indexed facet. The other four keys select matching CVRs before
the facet operation counts their index memberships.
| Key | Shape | Source |
|---|---|---|
has_website |
boolean | an eligible current website in company_web_presence, backed by an accepted ownership decision and retained primary-page reference |
has_linkedin |
boolean | kind linkedin |
has_facebook |
boolean | kind facebook |
has_instagram |
boolean | kind instagram |
has_web_contact |
boolean | any web_company_contact value for the company |
There is no missing != zero case: a company either has a resolved signal of
that kind or it does not. Enrichment is explicit-shortlist only (CONTEXT.md),
so has_*: true selects within the enriched shortlist. has_website: false
also includes a company with only unverified historical directory observations.
Out of scope. The observed profile signals
(line_of_business, value_proposition, and the rest) stay out of the
Filterable field list for the reason the registry section gives: they are
confidence-ranked model observations, not predicates.
Bulk participants contract¶
Status: implemented (#569, parent #428). The surface is the granted view
company_participants, readable by authenticated over PostgREST. It answers
leadership at scale: one row per current registry role across the whole
population, so a consumer fetches the participants of a page of
search_companies results with one cvr=in.(...) filter instead of one
company_detail call per CVR.
Columns: cvr, person_id, name, kind, relationship_type, role_name,
role_function, ownership_pct, voting_pct, valid_from.
ownership_pct and voting_pct are fractions in [0,1], despite their
_pct suffix: 1.0 means 100%, 0.25 means 25%, and 0 means a stated zero.
Null means no share was stated. Multiply by 100 only for percentage display.
These are source values, not values already scaled to a 0–100 percentage.
person_idis the registryenhedsNummerof the participant — stable across role changes.kindisperson,company, orother, so a company shown as legal owner reads like any other participant.relationship_typeis the edge type the registry states:legal-owner-of,beneficial-owner-of,has-voting-rights-in, andholds-role. The qualifiers are the source's own statement about the assertion, not about either entity.- Current roles only.
valid_to is nullis the view's own definition. An ended assertion is history:company_detail'srelationship_historycarries it with its period, source, and evidence. A past role never appears here. nameandkindcan be null: a participant removed from the identity projection keeps their edges, so the relationship is current while the identity projection is gone. A null name is not an empty name, and a null kind is not the valueother.- The view is live. It carries no
data_version; every read reflects the data as of that moment. A consumer needing the same identified snapshot as the grid uses the search and Facet contracts, which bind it. - Per-role raw-snapshot provenance is not part of this surface.
company_detailis the per-company evidence path.
What one company selection returns¶
Freshness¶
Every source slice reports its own state under current_detail.freshness:
the source, when collection last wrote it, and present or absent. That is
what makes an empty section readable — without it a consumer cannot tell
"collection has none" from "this slice is unavailable", and would be free to
read either as "the company has none".
source_history carries every effective-dated fact version the source stated,
each with valid_from, valid_to, is_current, the source update and
collection times, and the raw snapshot it was read from. Both parts are always
arrays: an empty one means the company has no stored history, which is a
different statement from the slice being unavailable. An unknown CVR sets
found to false and current_detail to null rather than returning an empty
object that reads like a company with no facts.
Stored fact types for a company today: name, status, company_form,
main_industry, secondary_industry_1.._3, address, postal_address,
phone, fax, email, website, lifecycle, registered_capital,
advertising_protected, and employment_annual / _quarterly / _monthly.
A production unit stores the same families minus the ones only a legal entity
has, plus parent_cvr.
Financials¶
current_detail.financials states its own window: window_years and
period_end_cutoff, then the retained publications newest accounting period
first. The cutoff is part of the payload because without it an empty list
reads as "this company files nothing" rather than "we do not retain that
far back" — and neither means zero.
Every publication carries its accounting period, correction state, document
references, promoted metrics, extractor and taxonomy version, and the absence
reasons from #105. Superseded publications stay visible with
is_current_for_period: false; a correction is a fact about the filing
history, not something to filter away.
metrics.ebitda is derived as metrics.ebit + metrics.depreciation, never
stored. It is null when either disclosed input is null.
absence_reasons is either a per-metric object or a single string when the
whole publication explains itself — "no_xbrl_document" for a PDF-only
filing, "not_parsed_yet" when XBRL exists but no metrics row has been
written. A publication outside the five-year window has no row at all, so it
is simply absent; the stated cutoff is what makes that absence readable.
Web signals¶
current_detail.web groups the observed values — presence, contacts, and
profile_signals — apart from the registry facts beside them. A CVR email is
something the company registered; a website email is something we read off a
page. The payload shows both and merges neither.
profile_signals is the current best-supported observation per field:
line_of_business, products_or_services, value_proposition,
target_market, commercial_posture, pricing_evidence,
customer_segments, geography. Each carries its source URL, page fetch
time, extractor and prompt version, model, confidence, and a bounded evidence
excerpt.
Two rules a consumer must not misread:
- A field with no row was not observed or not extracted. It never means the company lacks it.
confidenceis the extractor's own self-report, not a calibrated probability. The current view ranks by it — newest extractor version first, confidence only breaking ties within a version — and nothing may present it as a likelihood.
Earlier observations are never deleted: web_company_profile_signal keeps
every version's reading of a page so two can be compared against the same
evidence.
Entity mentions and discovered relationships¶
web.entity_mentions holds the companies that retained website or AI Mode
evidence names as a customer, supplier, partner, or competitor. A mention is
not a relationship: the source gave a name, not a CVR number, and those are
different claims.
A mention becomes an observed edge in relationship_history only after a
deterministic unique match — one normalized name matching exactly one
active company in the hub (#842: an ended company no longer competes). Zero
matches or several leave it a mention, with
resolution_state: "unresolved". Confidence plays no part in resolution: the
model's self-report is about whether it read the page correctly, not about
which legal entity a name refers to.
Resolution is reversible. Reversing it ends the edge rather than deleting
it, and the mention reads "reversed" rather than reverting to
"unresolved" — the assertion was made, and undoing it does not pretend
otherwise. The evidence that supported it stays on both.
Each edge also carries its own qualifiers — role_name, role_function,
ownership_pct and voting_pct, as the source stated them. Those belong to
the assertion, not to either entity, so a participant who is both a director
and an owner of the same company has two edges with their own. They were one
attributes object until #255; the object repeated its key names on every row
and stored an explicit null to say the source stated nothing, which cost 78 of
the table's 397 bytes a row for 289 distinct values across 1.2M rows.
Company-to-company ownership comes from Vrvirksomhed.deltagerRelation, not
from the participant surface — that index holds natural persons only. A
company participant carries no CVR number, so it is resolved through
enhedsNummer, a registry key that is unique in companies. An owner the hub
has not loaded yet produces no edge; the next reconcile re-reads the same
document and resolves it then.
Registry, observed, and derived edges keep their source_class in the
response for their whole life. A web-source claim and a register's statement
are never presented as the same kind of thing.
Website people¶
current_detail.web.people holds every person mention found on the company's
website. An empty list means that collection has no stored mention for the
company. It is never a missing field and never an error.
Each mention has the name, title, and company role as the website stated them. It also has its evidence, confidence, source URL, page fetch time, raw artifact reference, extractor version, prompt version, and model. An attributed work contact point stays under that mention. Each contact has its own evidence, source URL, page fetch time, and extractor provenance.
source_class: "observed" identifies the website claim. resolution_state is
unlinked, resolved, or reversed. An unlinked mention stays in the list
with equal standing. A resolved mention has a registry_person object with the
matched registry identity. A reversed mention stays visible, but its
registry_person is null because the link is no longer current.
Registry participants stay in current_detail.participants. Website mentions
stay in current_detail.web.people. The operation does not merge the two
source classes. It does not expose person_mentions,
person_mention_contact, person_contact_suppression, extractor results, or
LLM operation records as separate read contracts.
Participants¶
current_detail.participants holds the people and entities holding a role
now, each with their identity, kind, current name, and roles. An ended
role is not there; it is an ended edge in relationship_history.
Role assertions use distinct edge types, because the source states distinct
claims: legal-owner-of (the REGISTER register), beneficial-owner-of
(HVIDVASK), has-voting-rights-in, and holds-role for everything else.
Ownership and voting are separate edges — a share class can carry one without
the other.
ownership_pct and voting_pct use fractions in [0,1] throughout participant
roles and relationship history. For example, 1.0 is 100% and 0.25 is 25%;
null is unknown, not zero. The mapper parses the CVR source strings without
rescaling. The semantic quality rule warns on an out-of-range share but does
not clip, convert, or discard the source assertion.
tests/test_semantic_rules.py::test_ownership_and_voting_are_fractions_between_zero_and_one
checks the range and warning behavior.
tests/test_cvr_mappers.py::test_maps_fractional_source_share_without_converting_it
checks that the source fraction is not rescaled.
flowchart LR
CVR["CVR share string: 1.0"] --> Mapper["Parse numeric fraction"]
Mapper --> Ledger["Ownership or voting edge: 1.0"]
Mapper --> Quality["Warn when outside [0,1]"]
Ledger --> Reads["Company and person reads: 1.0"]
Reads --> Display["Consumer percentage display: 100%"]
ADR-0006 bounds this slice: a natural person contributes identity, kind,
current name, and their business roles. Their own stored history is the
participant kind and nothing else, because historical names are exactly what
the ADR excludes. people_roles is gone; company_participants is a view
over the relationship ledger, so there is no second copy of the truth.
current_detail.production_units holds the units the company operates now.
A unit it operated in the past is not there; it is an ended
operates-production-unit edge in relationship_history. That separation is
the point of the model: production_units.cvr is a projection of the current
edge, not separate relationship truth, so a P-number reassigned to another
company moves the column while both companies keep their edge.
Two rules worth stating plainly:
- A
hemmeligcontact point is never stored. CVR marks contact values the company asked not to have published. They do not appear in the current projection, in history, or in a retained artifact. - Only the newest employment observation per series is stored. The monthly
series runs to dozens of entries per company; the full series stays in the
raw snapshot.
employees_bandreads the annual series, which real companies leave empty —latest_employment_*is the one to use.
Graph-neighbor contract¶
Status: implemented (#572, parent #428). The PostgREST operation is:
company_graph_neighbors(
p_cvr bigint[],
p_min_confidence numeric default 50,
p_cursor jsonb default null
) returns jsonb
Only authenticated can execute it. The function is stable,
security definer, and has the fixed search path public, pg_temp. It is a
granted read over the relationship ledger — the graph is the stored current
observed edges, not a projection, and there is no publisher job.
It answers competitor traversal for the sales product: one or more CVR anchors in, neighboring companies out. The product proposes competitor candidates from this contract and does not infer new competitor edges.
Traversal rules:
- Types and classes. Only current (
valid_to is null)observededges ofcompetitor-ofandcustomer-ofare traversed. An ended assertion is history —company_detail'srelationship_history— andregistryedges are never neighbors, even if a mapper ever wrote such a type. - Symmetric. An edge observed in either direction makes each company the
other's neighbor. Each supporting edge keeps its own observation direction:
anchor-is-subjectreads "the anchor's page named the neighbor";anchor-is-objectreads "the neighbor's page named the anchor". - One row per neighbor CVR. With multiple anchors, the neighborhoods
union: a company connected to several anchors is one row whose
supporting_edgescarries every supporting anchor edge. Anchors themselves are never candidates.
Each item contains neighbor_cvr, neighbor_name, competitor_confidence,
and supporting_edges. Every supporting edge carries anchor_cvr,
direction, relationship_type, confidence (0-100), observed_at,
source_url, excerpt, and extractor_version.
Confidence and threshold rules:
competitor_confidenceis round(max edge confidence x 100) over the current observed edges between the neighbor and the anchors — the highest-support rule the consumer ranks by. No combined or averaged score exists in the contract, and the value stays the extractor's self-report, not a calibrated probability.- The threshold selects neighbors, never evidence. The default 50
(a null argument means the default) keeps a neighbor when its best current
edge reaches it;
supporting_edgesthen shows every current edge between the pair, including edges below the threshold. - Sorted by
competitor_confidencedesc, then neighbor CVR asc, in fixed 100-row keyset pages. The opaque cursor binds anchors, threshold, and position; a cursor for a different request is an error.
Error rules:
- An empty anchor array is an invalid-parameter error.
- An anchor that does not resolve to a company node is an invalid-parameter
error. The consumer resolves anchors to legal-company CVRs through
search_companiesfirst, so a non-resolving anchor is caller error, not an empty result. - A minimum confidence outside 0-100 is an invalid-parameter error.
The operation is live: it carries no data_version, the same posture as
company_participants. A page fetched after a web-enrichment run may mix
versions, and the cursor never expires from data changes. Snapshot identity
remains a search and Facet concern.
What was removed¶
company_registry_detail is gone. It was introduced in #215 to publish the
registry facts the mapper had started promoting, before the operation carried
them; it now returned a strict subset of current_detail, and two surfaces
for one question is how they drift.
No version-suffixed compatibility view was added in its place. Consumers and collection change together, which is the posture recorded under Change posture.
The population-filter and financial views stay — active_companies,
active_companies_financial_snapshot, financial_key_figures — because they
answer questions across companies, which a per-CVR operation cannot.
The population-filter and financial views stay — active_companies,
active_companies_financial_snapshot, financial_key_figures — because they
answer questions across companies, which a per-CVR operation cannot.
Hybrid-similarity contract¶
Status: implemented (#587, spec #584, ADR-0018). The PostgREST operation is:
hybrid_similar_companies(
p_seed_cvr bigint[],
p_filters jsonb default null,
p_cursor jsonb default null
)
security definer, stable, with a pinned search path, granted to
authenticated and revoked from public. The sales application calls it
when a user opens, refreshes, or pages the Prospect queue (consumer
ADR-0005); it sends approved Seed CVRs and its active policy's filters, and
applies tenant decisions to the returned page.
Ranking is Reciprocal Rank Fusion (k = 60, algorithm version hybrid-1.1)
over two lists:
- Semantic: cosine distance between the candidate's vector and each seed's vector at the contract's embedding version — the version most companies carry. Each candidate takes its best seed; other seeds never lower a result. A seed without a current vector still anchors the structured list.
- Structured: industry group (the first two digits of the registry branch code), employee band proximity (ordered employee groups within one), and revenue band proximity (DKK bands within one), counted against the best seed.
Each row carries the candidate CVR and name, the global hybrid_rank,
semantic_similarity (1 − best distance, 6 decimals), best_seed_cvr,
the structured matches that fired, the profile-signal evidence
references, profile_version, and signal_freshness (the newest page
fetch behind the candidate's profile).
Hard filters apply before ranking — a filtered-out candidate is invisible, not down-ranked:
| Filter | Meaning |
|---|---|
industry_group_in, industry_group_ex |
branch-code group allow/deny lists |
employees_min, employees_max |
ordered employee groups, 1–5: zero, 1–9, 10–49, 50–199, 200+; null or unknown source codes are invisible under the filter |
revenue_min, revenue_max |
DKK from the latest filing; a company without a filing is invisible under the filter |
exclude_cvr |
explicit exclusions |
municipality_ex |
JSON integer array, each code 1–9999, matched against the current registry municipality code; missing municipality is not a match |
keyword_ex |
At most 100 strings, each 1–200 characters after trimming; literal substring exclusions over current company name and profile text |
Municipality and keyword exclusions¶
accepted_prospect_evidence_contract() publishes both rules, their evidence
attributes, bounds, missing-data rules and cursor_filter_binding=exact_json.
The forward migration is
20260927100000_hybrid_policy_exclusions.sql.
Municipality is companies.municipality_code, the current CVR registry code,
not a postal code, name or inferred location. The numeric domain validates the
code representation; it is not a claim that every integer is an assigned code.
Keywords examine companies.name and each current company_profile_signals.value
for line_of_business, products_or_services, value_proposition,
target_market, commercial_posture, pricing_evidence, customer_segments
and geography. Each value is matched separately. URLs, evidence excerpts,
historical observations and vectors are not keyword text. Missing values and
profile text restricted by the person-evidence policy are not matches. A name
can still match when profile text is missing or restricted.
Trim ASCII spaces, tabs, CR and LF from each keyword, normalize both keyword
and source text to Unicode NFC, and apply PostgreSQL lower using the database
locale. Match the resulting literal substring with strpos. Accents and
internal whitespace remain significant. %, _ and regular-expression
characters are literal; there is no stemming or word-boundary rule. See the
PostgreSQL string functions.
A blank keyword, wrong JSON type or bound violation returns 22023.
Any matching keyword or municipality excludes the company. Empty arrays do
nothing. All hard filters apply together before both ranking lists, page
selection and accepted-card lookup. An excluded company cannot reappear on a
later page under the same current evidence. The accepted card reports
outside_context with no rank for the same policy. Available cards include
name and nullable municipality_code in their registry evidence, with source
times; their current profile observations already carry values and evidence.
The consumer reads these functions and has no new table grants.
The cursor binds the exact filter JSON. Changing an array's order, letter case or spacing requires a new first page, even when matching would be equivalent. Object key order is immaterial to JSONB equality. The read is live, so changes to source evidence can change results between pages; the cursor is not a snapshot.
An unknown filter key, a reversed range, an empty seed array, an unknown
seed CVR, and a request whose seeds carry no usable Semantic company
profile at the contract's embedding version are invalid-parameter errors
(22023) — the last is the consumer's own market activation gate, and
there is no structured-only fallback.
Candidates without a current vector stay outside the ranking. The response
counts them in pending_embedding_count, an unfiltered backlog count — companies with profile signals
whose embedding is not at the contract's version yet. They join the queue
when the pipeline stage or the backfill stores their vector.
The operation is live: no data_version, the same posture as
company_graph_neighbors. The keyset cursor binds the canonical seeds,
the filters, the algorithm version, and the position; a continuation must
repeat the same seeds and filters. Ranks are global, so page two continues
the numbering.
Implemented view surface¶
The following views are the current implementation from issues #26, #58, and
152. They stay accurate until the replacement migration lands.¶
| View | Base table(s) | Filter baked in |
|---|---|---|
active_companies |
companies |
ended_at is null (active only) |
financial_reports_latest |
financial_reports |
newest publication per (cvr, period_start, period_end), correction-resolved |
active_companies_financial_snapshot |
active_companies, financial_reports, financial_metrics |
active + at least one parsed report; that company's single most recent accounting period |
company_web_presence |
web_company_presence, companies, website evidence bundles/pages |
current best eligible signal per (cvr, kind) |
company_web_contact |
web_company_contact, companies |
every distinct contact value per (cvr, kind) |
active_companies columns: cvr, name, main_industry,
municipality_code, postal_code, employees_band, started_at,
municipality_name, main_industry_text. The industry text is the source
Danish label, or null when the source has none.
active_companies_financial_snapshot columns: cvr, name,
main_industry, municipality_code, postal_code, employees_band,
financial_period_end, gross_profit, revenue, result, equity,
ebit, ebitda, assets, profit_before_tax, staff_costs,
depreciation, current_assets, shortterm_liabilities, cash,
employees. A company absent from this view has no parsed nøgletal at
all (no XBRL, outside the backfill window, or not yet delta-fetched) —
never assume absence means zero.
ebitda here is derived (ebit + depreciation), not stored — see
financial_key_figures below. employees is the exact average headcount
the company reported in its filing, present in about half of filings;
employees_band beside it is CVR's band and is always present. They are
different things and neither substitutes for the other.
financial_key_figures columns: cvr, load_id, period_start,
period_end, ebitda, equity_ratio (soliditetsgrad), current_ratio
(likviditetsgrad 1), return_on_assets (afkastningsgrad),
return_on_equity (egenkapitalens forrentning), capacity_ratio
(kapacitetsgrad), gross_margin (dækningsgrad), operating_margin
(overskudsgrad). One row per parsed report, not per company — join to
financial_reports for the period, or use the snapshot view for "latest
per company". Every formula reproduces proff.dk's published figures for the
same filing. A null means an input was not disclosed, never zero, so
gross_margin and operating_margin are null for most companies because
turnover is not disclosed (ADR-0007).
Filing history is bounded to five accounting years (#190, ADR-0007's
2026-08-12 amendment). financial_reports retains only publications whose
period_end falls inside the window — the same predicate the nøgletal
backfill selects on — so a company's filings from before the cutoff are
absent from every view above, not merely unparsed. Absence therefore means
"outside the retained window" as well as "not filed digitally"; it never
means zero. Run snapshots contain only records selected by the five-year source
filter. Fetched XBRL documents in that window stay in GCS. The upstream
offentliggoerelser index is credential-free and unbounded, so widening the
window is a new metadata and document backfill with an earlier cutoff.
Why a financial metric is missing¶
A null metric used to conflate six different situations. Each one now has a
name, reported per metric by the operator-facing financial_metric_coverage
view (#105). Only the last two mean the extractor has a problem.
| Reason | Meaning |
|---|---|
no_xbrl_document |
The publication is PDF-only (~37% of filings). Nothing was parsed. |
not_parsed_yet |
An XBRL document exists, but no metrics row has been written for it yet. |
fact_absent |
The filing was read, and it does not disclose this fact. Danish ÅRL lets the smallest reporting class omit most of the income statement. |
ambiguous_context |
More than one whole-enterprise value for this concept and period. The extractor refuses to guess; a rule is missing. |
parse_failed |
The XBRL document could not be read at all. |
reason_not_recorded |
The row was parsed before reasons existed. Reasons are not backfilled. |
A publication outside the five-year window (ADR-0007) has no
financial_reports row at all, so it is absent from the view rather than
carrying a reason. None of these ever means zero.
"Region" maps to municipality_code (kommunekode), not postal_code —
CVR has no field literally named "region"; municipality is the coarser
filter the charter's M1 phrasing implies. Consumers needing finer geography
use postal_code, also exposed on the view.
Financial-band filtering uses gross_profit, not revenue — a
deliberate deviation from the PRD/charter's literal "revenue band" phrasing.
54's taxonomy spike found that Danish ÅRL lets the smallest reporting class¶
(Klasse B/micro — most of the corpus) omit turnover entirely from their
published statement, so filtering on revenue would silently exclude most
companies. gross_profit is the more universally present top-line figure;
revenue is still exposed for consumers who specifically want it (nullable,
same "missing != zero" rule).
Enforcement¶
Every base table (companies, production_units, people,
ingestion_runs, financial_reports, financial_metrics) has row-level
security enabled with no policies, which blocks any non-owner role
regardless of table-level grants — the ingester's own connection (table
owner) is unaffected. Only the implemented views are granted to the
authenticated role. A consumer with an authenticated
Supabase session can query a granted view; querying a base table directly
returns zero rows regardless of what it's granted, by construction.
This posture is preserved automatically across every companies/
financial_reports backfill and reconcile run (ingesters/cvr/loader.py's
bulk_load_rows): the staging+swap those runs use doesn't natively carry a
table's views, grants, or RLS state onto the new relation (Postgres binds
all three by OID, not name, and LIKE ... INCLUDING ALL doesn't cover RLS
or views at all) — the loader captures and re-establishes all three as part
of every swap. financial_metrics is a sibling table with an FK to
financial_reports.load_id that never itself participates in the swap;
Regnskab's reconcile additionally deletes any financial_metrics row for a
report the new scroll no longer contains, before the swap runs
(pipeline._delete_orphaned_metrics, _run's pre_swap_hook) — otherwise
bulk_load_rows re-adding the FK constraint after the swap would fail on
the orphaned reference.
Change posture¶
- Read contracts are released with their consumers. A replacement migration changes the current contract in place and removes obsolete surfaces. The project does not retain duplicate versioned views for backward compatibility.
- A contract replacement is not complete until collection migrations, tests, documentation, and known consumers agree on the same shape. The Known consumer register says which conformance suites must go green for a given surface.
- PostgREST exposes whatever the
publicschema grants allow — only the views are granted, so only they are queryable via the API regardless of PostgREST's own configuration. - Base-table columns and internal joins are not a consumer contract.
Performance (M1 acceptance criterion)¶
The M1-style query — active companies by industry × region × employee band
— is the read contract's validation gate (PROJECT_CHARTER.md's M1
criterion, company-filter half).
select cvr, name from active_companies
where main_industry = %s and municipality_code = %s and employees_band = %s;
Backed by a composite partial index:
create index companies_m1_filter_idx
on companies (main_industry, municipality_code, employees_band)
where ended_at is null;
Measured (tests/test_read_contract_performance.py, opt-in via
RUN_PERFORMANCE_TESTS=1 — not part of the routine suite, since generating
the population takes real time): 2,260,000 synthetic rows (procedurally
generated locally — no live upstream call was made or is needed for this),
matching companies' ADR-0005 population estimate. Query plan: Index Scan
on the composite index, 0.25ms execution time, 1.6ms measured
round-trip including Python/network overhead — comfortably under the <1s
target.
Performance (financial-band filter, #58)¶
select cvr, name from active_companies_financial_snapshot
where main_industry = %s and municipality_code = %s
and gross_profit between %s and %s;
Supported by companies_m1_filter_idx (company-side selectivity) plus:
create index financial_metrics_gross_profit_idx
on financial_metrics (gross_profit)
where gross_profit is not null;
The view is defined as a lateral join, not a plain distinct on over
the whole financial_reports/financial_metrics join — that was the first
attempt, and it measured 3.7 seconds at scale, over the <1s target: a
distinct on forces Postgres to resolve "most recent period per company"
across the entire population before an outer where on industry/
municipality/gross_profit can apply at all. lateral instead runs the
per-company "most recent period" lookup only for companies already matching
the outer industry/municipality predicate — Postgres inlines the view and
pushes that predicate down to companies_m1_filter_idx first, the same way
it already does for active_companies alone, so the lateral subquery only
ever executes for a small, pre-filtered set of companies. Kept here as a
record of what was tried and rejected, not just what shipped.
Measured (tests/test_regnskab_read_contract_performance.py, same
opt-in RUN_PERFORMANCE_TESTS=1 pattern): 2,260,000 synthetic companies +
1,000,000 synthetic financial reports/metrics (procedurally generated
locally — no live upstream call was made or is needed for this), yielding
900,538 companies with at least one parsed report. Query plan: Bitmap
Index Scan on companies_main_industry_municipality_code_employees_idx
narrows to the matching companies first, then a Nested Loop runs the
lateral lookup via Index Scan on
financial_reports_cvr_period_start_period_end_idx and
financial_metrics_pkey for just that narrow set — 0.4ms execution
time, 3.1ms measured round-trip including Python/network overhead —
comfortably under the <1s target, and consistent with the M1-only filter's
own measured result above.
Honest scope gap vs. the charter's exact M1 wording: PROJECT_CHARTER.md
asks for "positive gross-profit trend over 3 years" — a multi-year
comparison. active_companies_financial_snapshot exposes each company's
single most recent accounting period only, not a trend across periods; a
"positive trend" query would need to compare multiple financial_reports
rows per company (or a dedicated trend view) and isn't built here. This
issue's own acceptance criteria describe a single "revenue/gross-profit
band" filter, which is what's delivered — the trend clause is a residual
gap against the charter's original phrasing, not silently claimed as done.
Known cosmetic quirk: Postgres's LIKE ... INCLUDING INDEXES (used by
every staging+swap) regenerates index names from their columns rather than
preserving the name a migration gave them — companies_m1_filter_idx
becomes companies_main_industry_municipality_code_employees_idx after the
first swap. The index itself (columns, partial predicate, and the query
plan choosing it) is unaffected; only the name changes. Already tolerated
by #20's own index-preservation test. Not fixed here: correlating pre- and
post-swap index identity by definition rather than name is a real
generalization (analogous to the FK/view fixes already in bulk_load_rows)
but adds complexity for a cosmetic-only concern — worth revisiting if a
future need (e.g. pg_stat_user_indexes monitoring by name) makes the name
load-bearing.
Performance (widened filters, #630)¶
search_companies with filters and no text query walks companies in
lower(name), cvr order and stops at the first 101 that pass. The predicate
that must reduce the walked set is the registry filter. Layer 1 routed every
registry key through one company_search_matches_registry_filters(...) call;
that function does not inline (sub-selects in its body), so the four
set-shaped keys — company_form_codes, postal_codes,
secondary_industries, latest_employment_bands — could not drive the
partial indexes created for them, and a selective filtered browse scanned
every active company (66 s at 720k rows in testing).
630 applies those four inline in search_companies and¶
company_facet_counts as col = any(array(select …)) / && clauses. Under
the plpgsql custom plan the jsonb_array_length(…) = 0 guard folds away and
the index is used; company_search_matches_registry_filters (now the
residual range and boolean keys only) and
company_search_matches_cvr_keyed_filters (financial and web-presence,
cost 1000 so it sorts last in the filter) run only on the index-narrowed
rows. The financial keys keep the #58 LATERAL — "latest accounting period per
company", executed for the narrowed set, not the population.
Measured (tests/test_company_search_filter_performance.py, opt-in
RUN_PERFORMANCE_TESTS=1): 2,260,000 synthetic companies + 1,000,000
synthetic financial reports/metrics (procedurally generated locally — no
live upstream call). Filters: company_form_codes + one postal_codes
value + gross_profit band + equity_ratio floor + latest_filing_period
floor. 110 ms execution time, 170 ms measured round-trip including
Python/network overhead — comfortably under the <1 s target and consistent
with #58's own result. company_facet_counts on the same request stays
under the same bound.
For a request with only a financial range or has_parsed_financials,
company_facet_counts computes the latest financial row and the parsed-CVR set
once. It then semi-joins those CVRs to the active-company scan. It does not run
the latest-period LATERAL once per company. Facet selections apply to the
indexed dimensions after this financial constraint. Text requests evaluate
the financial predicate on their indexed search candidates.
The population test records the filter-only time separately and requires it to
stay below 10 seconds for 2,260,000 companies and 1,000,000 reports (#634).
Web enrichment (#152)¶
These views expose current presence and company-contact observations. Profile signals and entity mentions have separate read surfaces.
company_web_presence columns: cvr, name, kind, url, confidence,
evidence, step, extractor_version, retrieved_at. One row per CVR and
signal kind (website, linkedin, facebook, instagram) — the current
best signal, not the history.
company_web_contact columns: cvr, name, kind, value, source_url,
extractor_version, page_fetched_at. One row per CVR, kind, and value: a
company legitimately publishes several phone numbers, and each is its own
signal. Personal contact values never reach this view — they are dropped
inside the extractor under ADR-0006.
Which row wins¶
A website is eligible only when its presence row matches the evidence bundle's company, run, primary URL, and verified domain. The bundle must record an accepted ownership decision from a supported lead source, must not mark the page as a listing or parked site, and must retain a successful primary-page reference with its SHA-256 digest. An accepted related or unclear judgment remains accepted under the existing ownership policy; this read does not make a new ownership decision.
An old high-confidence llm-verdict row without that proof remains historical
evidence, not a current first-party website (#707). Among eligible rows,
confidence decides first, retrieval time second, and row ID breaks ties.
Non-website signal selection is unchanged.
Company detail, company search, has_website filtering, and website facet
counts use this same selector. The migration corrects affected indexed facet
memberships and advances the search publication revision. Pre-cutover cursors
expire. It does not delete or rewrite observations, bundles, or page artifacts.
tests/test_company_website_selection.py covers accepted versus unverified
history, rejected/listing/parked evidence, accepted related sites, and unchanged
social signals. test_company_search_contract.py::test_search_and_facets_exclude_unverified_directory_websites
covers the search and facet boundaries.
For the two historical records in #707, run
uv run python scripts/audit_website_presence.py. By default it reads the test
database secret and checks the retained GCS primary objects. It uses a read-only
database connection, bounded records and object sizes, and prints no page bodies
or contact values. DATABASE_URL can select another authorized audit database.
The output reports the actual current selection and retained history, not a
simulated policy. Version-1 evidence without a recorded object generation can
prove a current byte-hash match, but cannot prove the original generation.
Absence means unresolved, not absent from the web¶
A company missing from company_web_presence has no eligible current signal.
It may have no website, an unreachable site, no collection request, or only
unverified historical observations. The outcome and reason of each run remain
internal (company_enrichment_run), not part of this contract.
Operation tables stay internal¶
company_enrichment_run, company_web_request, source_retrieval, and
the evidence bundle tables are how the pipeline runs and what it retained.
They are execution mechanics, not findings; publishing them would make run
logs and provider task identifiers a contract we would then have to keep.
Row-level security gives the client roles no rows.
Shape and measured performance¶
distinct on (cvr, kind) over the whole table, unlike the LATERAL shape
58 needed for financials. CONTEXT.md constrains paid and scraped¶
enrichment to explicit shortlists, so this reads thousands of rows rather
than the 2.3M-row population that made distinct on too slow there.
Measured on a synthetic shortlist history of 269,789 presence rows over
50,000 companies (tests/test_web_read_contract_performance.py, PG14):
a single company's lookup plans onto web_company_presence_current_idx
(cvr, kind, confidence desc, retrieved_at desc) and executes in
0.068 ms. The index is what keeps the sort local to one company instead
of the population.
This historical timing predates the #707 ownership-evidence checks. It is not a measurement of the revised selector.
The revised selector passed the same 50,000-company shortlist test on local
PostgreSQL 15, with 269,789 historical presence rows and one ownership-backed
website per company. The first measured lookup returned three signal kinds in
0.092 s, within the test's 1 s budget. A subsequent EXPLAIN ANALYZE
reported 0.559 ms execution, using the presence, bundle, and page indexes.
These are synthetic local measurements, not live-service latency guarantees.
Exact website evidence references¶
Website contacts, profile signals, company mentions, person mentions, and person
contacts expose evidence_locator through their existing read surfaces and
company_detail. It contains retrieval_id, artifact (URI, generation, digest,
media type), projection_version, source_url, kind, start, and end.
Ranges use Unicode character offsets, with an exclusive end, in the retained
page-text projection. A supporting-page claim cites that supporting page. The
primary page is not a default citation. Earlier observations without exact
references have a null locator; they are not eligible sealed extraction inputs.
These remain source-qualified observations. A search citation is not an acquired website page, and an unresolved company mention is not a verified relationship.
Search source signals¶
company_source_signals is readable by authenticated consumers. Each row is an
extractor's observation of a retained organic or AI Mode source, with cvr,
goal, value, self-reported confidence, exact evidence, source_kind,
retrieval_id, acquired_at, retrieved_at and extractor target.
evidence_locator identifies the immutable raw artifact, JSON pointer and text
range. source_url is the provider's result URL when present; it is empty when
not supplied. No URL is invented from a query identity.
The surface keeps observations from each extraction run. An absent goal means
no signal was published; it does not mean the company lacks that capability or
relationship. Read run_id and request progress to distinguish failed and empty
extractions. These claims are separate from page contacts, verified website
presence, registry facts and canonical company relationships. AI citations do
not prove that the cited page was acquired. The input retains full answers,
blocks, citations and provider metadata. Current contact objections are applied
again when a parsed source result is reused.
Company search dimensions¶
The following additional keys work in search_companies and in the exact
company_facet_counts.total under the same data_version (#615):
| Key | Input | Meaning |
|---|---|---|
production_unit_count |
Integer range, for example {"min":2} |
Distinct P-numbers with an active membership in the latest collected CVR penheder array. |
purpose_text |
Nonblank string, up to 500 characters | Case-insensitive literal substring of the current CVR FORMÅL text. % and _ are literal characters. |
address_changed_since |
ISO date, for example "2025-09-01" |
A change between two registered location addresses, effective on or after the date and no later than today. |
has_email |
Boolean | At least one collected web company email contact. |
has_phone |
Boolean | At least one collected web company phone contact. |
Different keys use AND. has_email and has_phone follow the existing
has_web_contact convention: false means no retained matching observation.
They do not inspect registry contacts or prove that no contact exists online.
Use has_registry_email and has_registry_phone for registry evidence.
The first known address is not an address change. Repeated identical address
periods do not count; an apartment or floor change does. Dates use UTC and
are inclusive. Future changes are excluded. address_changed_since rejects
invalid dates and relative date words.
A collected empty production unit list means zero. Missing source data and unreadable current membership dates remain unknown and do not match a range. Expired and future memberships do not count. A conflicting current purpose has no searchable value. New source fields are populated by a normal CVR refresh; this migration does not fabricate values for existing rows.
Production unit count, purpose text, and address changes are filter-only.
Email and phone expose exact boolean facet buckets. Company form, website,
and advertising protection also expose per-value counts. Each bucket removes
only its own filter. All other filters, the text query, and the result
data_version still apply.
Company structure and financial classes¶
ADR-0027 defines these collection-owned filters and their evidence rules:
| Key | Input | Public meaning |
|---|---|---|
is_holding |
true, false, or "unknown" |
Registered holding activity under the source's DB07 or DB25 industry contract. |
group_role |
Array of parent, subsidiary, independent, unknown |
Observed majority voting control. independent means “No registered majority-control link” and requires complete coverage. |
performance_class |
Array of high_growth, liquid, profitable_cash_rich, financially_challenged |
High gross-profit growth; Current assets cover short-term liabilities; Profitable with cash coverage; Negative equity or a loss with a current-asset shortfall. |
Values within a set use OR. Different keys use AND. Parent and subsidiary can both match one company. Financial classes can also overlap. All three keys have exact disjunctive facets. Holding and group facets retain an explicit unknown bucket. Financial facets count true matches only; unknown financial inputs never become a positive match.
Collection retains industry versions, voting intervals and source locators. An unbounded successful CVR population run plus complete company relation collections establishes group coverage. Missing organisation data, unresolved votes, contradictory controlling parents, and control cycles prevent an independence claim. Natural-person ownership does not create a company parent.
Financial extraction version 6 retains monetary facts with entity, context, currency, period, document hash and source locator. Incompatible or ambiguous facts remain unknown. The newest report revision takes precedence even when unreadable. Financial windows require consecutive annual periods and expire 18 calendar months after the latest eligible period. No currency conversion or estimate fills a gap. Re-extraction of the same document keeps its original fetch time.
company_detail.current_detail.classifications returns the rule version,
evaluation date, publication version, truth values, reasons, and source
evidence. Holding and control use precomputed effective intervals. Financial
classes use a stored projection with validity dates. Reads select the UTC date;
they do not traverse control graphs or rebuild financial histories. Source
writes and full population swaps refresh these projections in the same
transaction. The result version also binds the evaluation day.
Existing rows need a normal CVR refresh and financial re-extraction to acquire
these source inputs. Until then, unavailable evidence remains unknown.
The full-population performance proof passes
the 500 ms p95 limit with populated control and financial evidence (#770).
All six sales consumer conformance tests pass against this migration chain
(consumer PR #189, issue #190). Consumer CI still reads producer main and
must pass again after the producer contract is merged. These checks do not
constitute a deployment.
Test hub anonymous sign-in¶
On 2026-09-12 the test hub's /auth/v1/settings response reported
anonymous_enabled: true and disable_signup: false (#671). Anonymous
sign-in was already enabled, so this work did not change the setting. This
check did not create a user or test a browser sign-in.
An anonymous sign-in creates an authenticated session. It is distinct from an unauthenticated request with the public API key. Existing RLS and grants still govern reads; see the Supabase anonymous sign-in documentation.
Company-form vocabulary¶
company_form_vocabulary() is granted to authenticated and returns code
and nullable name, ordered lexically by code. It enumerates distinct non-null
company-form codes from retained CVR company records, including ended companies.
Names are the source-disclosed Danish short descriptions already collected with
those codes. The latest source-updated nonblank name wins; the lowest CVR breaks
an equal timestamp. If no record discloses a name, name is null.
This is the collected vocabulary, not a complete legal code catalog. The service
does not invent labels or infer that an absent code is legally retired. A code
remains listed while retained company records use it. After it disappears from
the dataset, a restored selection must retain its code and show its missing name.
company_form_codes continues to accept that selection; no current matches means
an empty result. Consumers read no base tables and keep no copied code catalog.
The vocabulary supplies names, not facet counts. company_facet_counts supplies
exact context-dependent counts under the search publication version. Missing or
pending counts must not be rendered as zero. The company-form vocabulary itself
is a current read and is not bound to a search cursor.
Person list and profile¶
Issues #768 and consumer #116/#161; ADR-0028.
person_read_contract() publishes the live filter definitions, response field
sets, page size/order, query bound, and identity/name/role/contact/claim states.
Search validates keys, set sizes, and role values from this operation. A
consumer conformance test must compare its handlers with this live vocabulary.
Both data operations are stable, security-definer reads with a pinned search path.
Only authenticated has execution access. Private tables and helper views
have no consumer grant.
select search_people(
p_query := 'hansen',
p_filters := '{"company_cvrs":[41527080],"relationship_types":["holds-role"]}',
p_cursor := null
);
select person_profile(p_person_id := 4000000001);
The list contains CVR PERSON participants only. person_id is the registry
enhedsNummer, unchanged across names, roles, companies, and publications.
Company participants and unknown kinds are excluded. Website mentions remain
company-scoped observations, not list identities. Unresolved mentions remain
available through company_detail; they are never silently merged here.
Search accepts a trimmed, case-insensitive literal substring of the current
registry name, or an exact decimal person ID. % and _ are literal text.
Null or blank means no text restriction. The limit is 200 characters. It does
not search old names, website names, contact values, or titles.
| Filter | Meaning |
|---|---|
company_cvrs |
1–100 eight-digit numeric CVRs; a current registry role at one of these companies. |
relationship_types |
1–100 values from holds-role, legal-owner-of, beneficial-owner-of, has-voting-rights-in. |
Values combine with OR; keys combine with AND on the same role. Duplicate
values and set order do not change the request. Unknown keys, invalid types,
empty sets, malformed cursors, and invalid values raise SQLSTATE 22023.
No filters means all known registry person identities, including those whose
roles have ended. These filters do not claim employment or decision authority.
The response has items, total_count, data_version, and next_cursor.
It sorts by ascending numeric person_id, with 100 items per page. Each item
has person_id, name, name_state, state, current_roles, and
work_contact_state. A blank or absent name is null with name_state=unavailable.
The contact state is available, withdrawn, or unavailable; list rows carry
no contact values. Multiple roles never duplicate a person row.
The ownership and voting shares in current_roles and profile role_history
use the same fractional unit as Participants:
ownership_pct = 1.0 or voting_pct = 1.0 means 100%. Null means no stated
share. Neither the person reads nor the mapper multiplies these values by 100.
Pass the opaque next_cursor unchanged with the same query and filters. It
binds the last person ID and a transactional publication version. Person,
role, identity, company, contact, suppression, and mention changes invalidate
it, as does the UTC date. Changes to other company publications may also
invalidate it. An expired cursor raises 22023 with person cursor expired;
restart at the first page. Totals and rows use one database statement snapshot.
Uncommitted changes are not visible. There is no historical snapshot service.
person_profile returns person_id, state, identity, data_version,
current_roles, role_history, website_claim_state, website_mentions,
and work_contacts. An unknown or non-person ID returns state=unavailable,
null identity, and empty evidence collections. If the current identity row was
removed but retained participant-kind evidence identifies a person, the ID
remains readable with state=withdrawn and no current name. History is retained.
The identity contains the current registry name, kind, source, collection time,
and field evidence for the name. The current-name projection does not retain
a source update timestamp: source_updated_at is explicitly null. This read
adds neither historical names nor residential data. participant_facts_as_of
already takes a person ID and date; it remains the narrow kind-history read.
Each role carries its company CVR, current company name or unavailable state,
type, source role/function, ownership/voting percentages, inclusive source
period, and source provenance. evidence_id is a stable hash of the source
assertion's natural key, not a confidence score. Correcting its values or end
date does not change that key. Different source records remain separate.
state is current, ended, future, or unknown_period. A missing start
cannot prove current membership. role_history includes every retained
registry assertion; current_roles contains only those current on the UTC date.
Every field of an assertion shares its source record and collection evidence.
Website mentions expose their own name/title/role and per-field evidence,
source URL, fetched/recorded times, extractor/prompt/model versions, and
self-reported confidence. They never replace registry fields. Distinct claims
within one company remain separate and give website_claim_state=conflicting.
Different roles at different companies are not a conflict. A reversed link
returns only mention_id and state=withdrawn, not the old person's claim.
work_contacts.items contains attributed work email/phone values permitted by
ADR-0008/0015, each with contact ID, mention ID, CVR, source evidence and state.
A suppression excludes a value before deletion. Deletion retains only contact
and mention IDs; the published item becomes {contact_id, mention_id, cvr,
state: "withdrawn"}. The aggregate is available if any permitted value remains,
withdrawn if only withdrawn references remain, otherwise unavailable. Missing
means not available in retained evidence, not proof of no real-world contact.
Deleting a whole mention also removes its contact references. Pre-migration
contact deletions cannot be reconstructed and remain unavailable.
Free-text person evidence and URLs may repeat contact values. Once a company has a suppression or retained withdrawal, those unstructured evidence fields are withheld for that company's person claims and contacts. Typed permitted values and source times remain readable. No suppression value, reason, note, raw personal artifact, or contact-entitlement state is returned.
Accepted-prospect evidence¶
Issue #769 and consumer #163; ADR-0028.
select accepted_prospect_evidence(
p_cvr := 41527080,
p_seed_cvr := array[12345678]::bigint[],
p_filters := '{"industry_group_in":["62"]}',
p_expected_versions := null,
p_linked_person_id := null
);
accepted_prospect_evidence_contract() publishes the current ranking versions,
accepted filter rules, seed bounds, response fields, missing-reason vocabulary,
and card/linked-person states. Queue and card validation share these definitions;
consumers can check their handlers against them without reading private tables.
This authenticated read directly selects one CVR after the global Hybrid calculation. It does not enumerate or page through the runtime queue. Use the original target-market seeds and hard filters. Do not pass the queue's current seller-state suppression list as that context: an explicitly excluded CVR remains outside the context. The tenant store keeps references and seller state; it does not copy platform evidence.
The result is evidence_mode=current, with read_at, cvr, state,
reason_code, context, current_versions, ranking, reason, and
linked_person. It does not reproduce the rank at acceptance. The ranking
and reason can change when source evidence or the candidate population changes.
The optional expected versions contain exactly algorithm and embedding
from a prior Hybrid response. They are policy expectations, not a snapshot ID.
| State | Meaning |
|---|---|
available |
The CVR currently ranks in this saved context. |
unavailable |
reason_code identifies company_unavailable, seed_unavailable, seed_profile_unavailable, profile_unavailable, or outside_context. Rank/reason are null. |
withdrawn |
A supplied prior version reference can no longer be evaluated because a required company, seed, or profile is missing. No prior evidence is replayed. |
policy_changed |
ranking_version_changed: supplied algorithm/embedding versions differ from the current versions. Rank/reason are null; the caller can explicitly request the current policy. |
The rank is the ordinal position in the saved context, never a percentage
or a sale probability. It shares hybrid-1.1, RRF with k=60, and the dominant
embedding version with hybrid_similar_companies. Hard filters apply before
ranking; the direct lookup applies after ranking. The response names the best
seed, semantic similarity, versions, and rank meaning. Queue and card reads
select tied financial periods by period start, publication date, then load ID,
so corrections and their evidence have a deterministic order.
The reason has code hybrid_similarity, a factual description of the ranking
basis, structured_matches, and supporting evidence: registry company
facts with source times, selected financial publications and revenue, stored
embedding versions and creation times, and current profile field observations
with URLs and fetch times. A stored vector's profile_version can differ from
the latest profile observations; the separate versions/times make this visible.
A missing optional fact is not zero or a negative finding. This reason is a
similarity basis, not a new model of willingness to buy or permission to contact.
A linked person is never inferred from registry participants. A caller may
supply a seller-selected p_linked_person_id. The read returns that registry
reference only if an unreversed mention for this CVR has a currently
permitted attributed contact. The object has state, person_id, and
mention_ids, with no contact values. No selection or unsupported evidence is
unavailable. Reversed or withdrawn supporting contact evidence gives withdrawn
with a null person ID. Several supported mentions can qualify the same selected
person. This does not choose a person, buy a contact, or grant an entitlement.
Seeds must contain 1–100 eight-digit CVRs. Filter keys match the existing Hybrid
contract; malformed sets/ranges and unknown keys raise 22023 before data
availability is checked. These validation rules also apply to the queue read.
The private shared ranking routine, person helpers, and base tables have no
consumer grant. The authenticated producer conformance suites cover direct
lookup beyond 100 candidates, shared ranks, source evidence, missing states,
changed policy, selected-person scope, and withdrawal.
Profile evidence values and URLs in accepted-card evidence are withheld after a company contact suppression or retained withdrawal, because these fields can repeat a removed contact. Their state is unavailable; source times and versions remain readable. The queue also withholds affected profile evidence URLs. The same private privacy predicate governs these reads and person reads.