Skip to content

Read contracts

This page separates the implemented read views from the accepted company-detail contract. Consumers, such as data-analysis, use these published surfaces. They never read base tables or call upstream sources.

Schema reference is the one index of every surface granted to authenticated, mapping each to its migration and to its section here. Known consumers records which repository depends on which surface. A consumer scopes a data feature by enumerating from those two pages and the migrations, not from the fields its application already reads (ADR-0019).

Company-detail contract

Status: implemented (#215, #216, #217, #218, #219, #220, #221, #507). The company-detail operation is the single company-selection surface.

The operation is the PostgREST function company_detail(p_cvr bigint). It is security definer with a pinned search path over base tables that keep row-level security with no policies, and stable, so it reads and never writes. company_facts_as_of(p_cvr bigint, p_as_of date) answers what the source stated on a given date.

Selecting one company by CVR returns one object with three required parts:

{
  "current_detail": {},
  "source_history": [],
  "relationship_history": []
}

current_detail contains the current company, production units, participants and roles, registry contact points, web signals, and financial publications and metrics. Financial publications stay limited to the latest five accounting years.

current_detail section Required content
company identity, names, lifecycle and status, legal form, primary and secondary industries, locations, advertising protection, registry contacts, employment, and current company attributes
production units identity, current parent, names, lifecycle and status, industries, locations, registry contacts, and employment
participants participant identity and kind, current allowed name, current roles, role functions, ownership, and voting rights
financials five-year cutoff, publications, correction state, document references, metrics, and absence reasons
web current presence, web contacts, company-profile signals, entity mentions, and website people, each with evidence
freshness source-slice state, source update time, collection time, and refresh state

source_history contains effective-dated source facts for the company and its related CVR entities. It is queryable without reading GCS. Each fact identifies the entity, source, fact type and value, effective period, source update time, collection time, and raw snapshot.

relationship_history contains current and ended edges that touch the company. Each edge identifies both entities, relationship type, effective period, source and trust class, observation time, and evidence. Registry edges and observed web edges remain distinguishable.

other_entity identifies the counterpart with kind (company, production-unit, or participant), source_key (CVR number, P-number, or enhedsNummer), and name: the counterpart's current registry name (company name, production-unit name, or participant name within ADR-0006). The name is the current name, not the name at the time of the edge, so an ended role or production unit is named too. It is null when the registry has no name. participant_kind is the participant's person, company, or other kind and is null for a company or production-unit counterpart (#870).

The response includes source-slice freshness and an explicit absence reason when data is missing. A missing value does not mean that the source published zero, false, or an empty relationship.

The public surfaces are:

Surface SQL name Status
company-detail operation company_detail(bigint) implemented, company slice only
company-search operation search_companies(text, jsonb, text, text, jsonb) implemented, fixed 100-row pages with Facet filters
company-Facet operation company_facet_counts(text, text, jsonb) implemented, exact disjunctive counts
company-search filter contract company_search_filter_contract() implemented, self-describing accepted filter keys
bulk participants view company_participants implemented, current registry roles only
graph-neighbor operation company_graph_neighbors(bigint[], numeric, jsonb) implemented, current observed competitor-of and customer-of edges
hybrid-similarity operation hybrid_similar_companies(bigint[], jsonb, jsonb) implemented, RRF over semantic and structured ranks at request time
as-of source read company_facts_as_of(bigint, date), production_unit_facts_as_of(bigint, date), participant_facts_as_of(bigint, date) implemented
current-company view active_companies implemented
current web-signal views company_web_presence, company_web_contact, company_profile_signals, company_entity_mentions implemented
current financial views financial_reports_latest, active_companies_financial_snapshot, financial_key_figures implemented

entity_nodes, entity_facts, and entity_relationships are not consumer surfaces. They are read through the operation, which is the only granted path to them.

Company-search contract

Status: implemented (#550, spec #549).

The PostgREST operation is:

search_companies(
  p_query text default null,
  p_filters jsonb default '{}'::jsonb,
  p_sort_by text default 'relevance',
  p_sort_direction text default 'asc',
  p_cursor jsonb default null
) returns jsonb

Only authenticated can execute it. The function is stable, security definer, and has the fixed search path public, pg_temp. Base tables stay behind row-level security.

The result has items, next_cursor, and data_version. items has at most 100 rows. Each row contains cvr, name, domain, main_industry, main_industry_text, municipality_code, municipality_name, status_group, postal_code, employees_band, started_at, and the latest financial_period_end, gross_profit, revenue, result, equity, ebit, derived ebitda, and assets.

municipality_name is the current registry municipality name for municipality_code, the same value as active_companies.municipality_name. It is null when the registry has no name. status_group is the token that the status_groups filter and company_facet_counts().facets.status_groups use for the company: active, bankrupt, voluntary_liquidation, compulsory_dissolution, reassumption, or unknown when the registry status is absent or not mapped. Both fields are present in filtered and unfiltered searches (#904); the authenticated regression test_company_search_contract.py::test_company_search_rows_return_municipality_name_and_status_group covers both paths and missing values.

main_industry_text is the source Danish industry label, including DB25 labels. It is present in both filtered and unfiltered searches and is null when the source has no label. Do not derive it from a bundled DB07 code table. The same field is available on active_companies. The authenticated regression test_company_search_contract.py::test_company_list_surfaces_return_source_industry_text covers both search paths, exact-CVR search, source labels, and missing labels.

The operation starts from every active company. For the existing registry sorts, it left-joins optional financial output after selecting the candidate page. Financial ordering reads the latest parsed figures before selecting the page. A company with no resolved website or no parsed financial report stays in the result with null values. It never receives zero as a substitute.

Search order is:

  1. exact CVR;
  2. exact case-insensitive legal name or normalized website domain;
  3. legal-name or domain prefix;
  4. pg_trgm fuzzy legal-name or domain match at the tested 0.3 threshold.

Without an effective filter, the existing sorts are relevance, cvr, name, main_industry, municipality_code, postal_code, employees_band, and started_at, in either direction. With an effective filter, the existing registry order is name ascending.

Both filtered and unfiltered searches also accept gross_profit, revenue, result, and employees with p_sort_direction = 'desc' (#859). The financial sorts use the same latest parsed accounting period as the returned financial figures: latest period end, then latest publication, then load ID. employees uses the latest numeric CVR employment count, not the text employment band or the average headcount from a filing.

The selected value is primary and CVR ascending is the final stable tie-breaker. Missing values sort last, including when the next page crosses into missing values. These are server-side keyset orders over all matching companies, not a sort of the first result page. An unsupported sort or direction is an invalid parameter error.

The filter object has industries, employee_groups, municipality_codes, and status_groups. Values in one Facet use OR. Different Facets use AND. Missing values use unknown; they never use zero. The cursor is a JSON object that the consumer treats as opaque. It binds the normalized query, normalized filter object, sort, last sort value, final CVR, and data_version. data_version combines the latest successful CVR company run, the latest successful Regnskabsdata publication run, and the latest terminal web-enrichment run, the financial mutation revision, the current report relation identity, and the UTC retention date. If one changes, the operation rejects the old cursor with company search cursor expired because data changed. The consumer starts at page one and does not mix daily versions.

tests/test_company_search_numeric_sorting.py exercises the authenticated numeric-sort contract: latest figures, registry headcount, empty filters, ties across pages, the transition to null values, and rejection of changed query, filter, sort, data version, or malformed cursor values.

Company-search performance

The first implementation materialized all active companies before fuzzy candidate selection. It measured 35.686 seconds on 2,260,000 active synthetic companies and was rejected.

The original registry-sort query uses separate indexed paths for exact CVR, exact name, prefix, and trigram candidates. It selects at most 101 CVRs before joining financial output. These measurements do not cover numeric financial ordering, which must inspect each matching company's latest figure. On the same full population, the original query measured:

Request Execution time
one-letter fuzzy company-name search 204 ms
default first page, sorted by name 41 ms

Both meet the 500 ms requirement. The reproducible opt-in test is tests/test_company_search_performance.py with RUN_PERFORMANCE_TESTS=1. The measurements use local synthetic data and are not a production capacity forecast.

Company-Facet contract

Status: implemented (#558, parent #428). The PostgREST operation is:

company_facet_counts(
  p_data_version text,
  p_query text default null,
  p_filters jsonb default '{}'::jsonb
) returns jsonb

Only authenticated can execute it. The result contains data_version, calculated_at, total, and exact Facet values for industry, employee group, municipality, status, company_form_codes, advertising_protected, has_website, has_email, and has_phone. Search applies before every count. Each Facet ignores its own selection and keeps the selections from all other Facets. A selected value with zero matches stays in the result. Boolean facets always return JSON true and false buckets, including zero counts. Other zero values are absent. Legal-form values are source codes with source labels (or the code if the label is absent); a missing form has no bucket. Website, email, and phone refer to the retained web-presence/contact rows, as their filters do. Multiple source rows for one CVR count once. Registry contact fields remain separate. A changed data version is an invalid parameter error.

total (#617) is the exact size of the whole matched set for the request — the number of companies search_companies would return for the same query and p_filters, with every filter applied (all Facet selections AND-combined, plus the registry / financial / web-presence Filterable fields). It is the count for range and trigger filters without per-value buckets: the consumer's "X virksomheder i målgruppen". With no status selection, sum(status_groups[].count) equals total under the same snapshot.

Industry values are exact six-digit leaf codes and source Danish text. Employee values are zero, 1_9, 10_49, 50_199, 200_plus, and unknown. Municipality values use the source code and the hub's official kommunekode to municipality-name reference; the source kommuneNavn is only used when a code is not in that reference. Status values are active, bankrupt, voluntary_liquidation, compulsory_dissolution, reassumption, and unknown. The active grid does not include ended companies.

All twelve facet groups use the complete company_facet_cube. The index stores exact counts for shared combinations and effective-date intervals. A CVR membership relation connects text-search and filter-only matches to those same combinations. Each request selects the current UTC interval, so holding changes, control changes, and financial expiry do not need a midnight publication job. The count query groups values before building JSON and uses at most 32 MB of working memory per query-plan operation.

Source statements update the affected memberships and counts in their transaction. Large source statements rebuild the index; an atomic company-table replacement publishes the control projection and index before commit. The CVR loader writes pages as set-based upserts and keeps the final record for a repeated CVR. An unchanged control publication does not cause a graph rebuild for each metadata update. Web-presence and contact writes also update the index and publication revision. A caller must repeat search if its data version is no longer current.

The opt-in proof uses 2.26 million synthetic companies and 20 measured samples after one warm-up per context. It includes source company forms and holding intervals, 6,000 companies with control evidence, 12,000 companies with four annual filings each, and web presence and contacts. Its cases cover broad, selective, fuzzy, presence, structure, financial-class, combined, and restored-form requests. It checks literal expected class counts and records row-at-a-time upserts, page upserts, and full replacement throughput. This controlled fixture is not a measured production distribution.

With the previous implementation, the original fixture reproduced an unfiltered p95 of 8,459.6 ms, four-facet p95 of 144.2 ms, and fuzzy-search p95 of 106.1 ms. The complete index passed all eight contexts on local PostgreSQL 17 with the populated fixture. These p95 values use 20 samples after one warm-up:

Context p95
Unfiltered 404.9 ms
Original four facets 426.3 ms
Fuzzy search and facets 122.4 ms
Website, email, and phone presence 434.9 ms
Holding and group roles 348.8 ms
Financial classes 331.9 ms
Combined facets 343.4 ms
Restored absent company form 328.7 ms

Source-write measurements include projection and index publication. Unchanged single-row upserts reached 370 rows/s; unchanged 1,000-row pages reached 2,247 rows/s. Pages that changed advertising-protection values reached 237 rows/s. A full replacement of 2,260,000 companies took 395.1 seconds (5,720 rows/s) and retained the exact public total. An explicit complete-index refresh took 355.6 seconds. Initial population and index creation took 214.6 seconds; control, financial, and web evidence publication took 103.7, 110.9, and 11.2 seconds. These are local measurements, not production throughput guarantees.

Run the proof on a local disposable database with all migrations applied:

DATABASE_URL=postgresql://localhost/disposable RUN_PERFORMANCE_TESTS=1 \
  uv run pytest tests/test_company_facet_performance.py -q -s

Self-describing surfaces

Status: implemented (#629, spec #626, ADR-0019). A published surface reports its own accepted shape from the same place it enforces it, so a consumer reads the contract without reading a base table and the reported shape cannot drift from the code.

company_search_filter_contract() is the filter contract for Company search. It is security definer, stable, fixed search path, granted to authenticated, revoked from public. It returns one row per accepted p_filters key:

Column Meaning
filter_key The p_filters key, e.g. industries.
value_domain The scalar type of each value: text, integer, numeric, date, or boolean.
filter_shape set — the value is a JSON array, matched any-of; range — the value is {"min": x, "max": y} (either bound optional, at least one required), the company value must be non-null and inside; boolean — the value is JSON true / false, matched against a presence flag.
vocabulary_kind set keys with a vocabulary: enum — a fixed list; reference — values from a named reference. Null when no vocabulary applies.
enum_values The fixed list, when vocabulary_kind is enum; null otherwise.
reference_name The reference the values come from (cvr_branchekode, kommunekode, company_form_vocabulary, dk_postnummer, cvr_ansatte_kode), when vocabulary_kind is reference; null otherwise.
accepts_unknown_sentinel Whether the key also accepts the string unknown for a company whose source value is missing.
facet_capable Whether company_facet_counts returns an exact disjunctive bucket list for this key. Range and trigger keys are filter-only.
read_contract_anchor The section of this page a consumer reads for the key's semantics.

The original four facet_capable keys — industries, employee_groups, municipality_codes, status_groups — are anchored to Company-Facet contract, are set-shaped, and all have accepts_unknown_sentinel true: a company whose source value is missing matches unknown, never zero. accepts_unknown_sentinel is stated per key so a consumer validating from this contract alone knows unknown is allowed even for a reference or integer key, where it is not in enum_values. The wider registry set is at Filterable fields: registry. company_form_codes, advertising_protected, has_website, has_email, and has_phone also report facet_capable = true; their accepted filter shapes and source meanings are unchanged.

The function assembles its answer from company_search_filterable_field, an immutable reference table. Adding a Filterable field is one row there; the contract and the enforcement below both follow from it.

Unknown-key and malformed-value rule. search_companies and company_facet_counts read the same reference table. A p_filters key that is not in it is an invalid-parameter error (22023) whose message names the key — a consumer typo or a stale key is an error, not a silent no-op. A p_filters value that is not a JSON object, a set value that is not an array, a boolean value that is not true / false, or a range value that is not a {min, max} object with at least one bound, is the same class of error. The key set the contract reports is exactly the key set the two operations enforce.

Company CVR set filter

cvrs restricts search_companies and company_facet_counts to an explicit set of active companies (#860). A consumer can save this filter with the other Saved view filters; the producer does not store a list.

{"cvrs": [10050715, 10000297], "municipality_codes": [101]}

The contract reports filter_shape = set, value_domain = integer, accepts_unknown_sentinel = false, and this section's anchor. Supply 1–5,000 eight-digit JSON integer CVRs. Strings, fractions, nulls, booleans, values outside 10000000–99999999, an empty array, and more than 5,000 entries raise SQLSTATE 22023; the message names cvrs. Duplicate CVRs do not duplicate companies. An inactive or unknown CVR contributes no row and does not cause an error.

The set uses OR within cvrs and AND with the query and every other filter. Exact-CVR p_query remains an intersection, not an override. The normal filtered name and numeric sort rules, 100-row pages, and bound cursors apply. CVR membership narrows candidates before the other predicates. Facet totals and disjunctive buckets remain exact within that set; there is no CVR bucket. tests/test_company_search_filter_contract.py covers the public contract, input bounds, intersections, paging, numeric ordering, and facet counts.

Filterable fields: registry

Status: implemented (#630, spec #626). Company search can filter on every registry attribute the collection sources. Each key is a row in company_search_filterable_field, enforced by search_companies, published by company_search_filter_contract(), and — for a range or boolean key — applied by company_facet_counts as a filter-only constraint on the counted population (the pre-aggregated cube is left for a company scan when such a key is present).

Key Shape Company attribute
company_form_codes set (company_form_vocabulary) company_form_code
postal_codes set (dk_postnummer) postal_code
postal_code range (integer) postal_code; composes with postal_codes using AND
secondary_industries set (cvr_branchekode) secondary_industries — array overlap, any-of
latest_employment_bands set (cvr_ansatte_kode) latest_employment_band
registered_capital range (numeric) registered_capital (DKK)
incorporation_date range (date) started_at — company age / incorporation date
latest_employment_count range (numeric) latest_employment_count — the exact latest headcount
advertising_protected boolean advertising_protected
has_registry_phone boolean a registry phone is present
has_registry_fax boolean a registry fax is present
has_registry_email boolean a registry email is present
has_registry_website boolean a registry website is present

missing != zero. A range filter on a nullable attribute (postal_code, registered_capital, started_at, latest_employment_count) excludes a company that has no value — it is never read as zero or as the epoch. A set filter likewise matches only a company whose attribute is one of the listed values; a null attribute is not a match. advertising_protected and the has_registry_* flags are always determinate.

Out of scope. Profile-signal semantic fields (line_of_business, products_or_services, value_proposition, and the rest of web signals) are not Filterable fields — they are model observations, not registry facts, and a filter would present a confidence- ranked reading as a hard predicate. Participant and graph-neighbour attributes (a person's role, an ownership percentage, a competitor edge) are answered by company_participants and company_graph_neighbors, not by a Company-search filter.

Filterable fields: financial

Status: implemented (#630, spec #626). Company search can filter on the financial statement the collection parses. Every figure and ratio is read from the company's single most recent parsed accounting period, correction-resolved — the LATERAL "latest per company" shape #58 measured and kept (see Performance (financial-band filter, #58)). Search applies company_search_matches_cvr_keyed_filters(cvr, filters) to candidates. Figure and ratio keys are filter-only. Facet requests first select matching CVRs, then use their index memberships to count the facet values. Requests without text resolve the latest financial rows as one set.

Key Shape Source
gross_profit range (numeric) financial_metrics.gross_profit (bruttofortjeneste)
revenue range (numeric) financial_metrics.revenue — disclosed by a minority of filings (ADR-0007)
equity range (numeric) financial_metrics.equity
result range (numeric) financial_metrics.result (årets resultat)
ebit range (numeric) financial_metrics.ebit
ebitda range (numeric) derived: ebit + depreciation, never stored
assets range (numeric) financial_metrics.assets
filing_employee_count range (numeric) financial_metrics.employees — the exact headcount the filing reported, not latest_employment_*
equity_ratio range (numeric) soliditetsgrad: equity / assets
current_ratio range (numeric) likviditetsgrad 1: current_assets / shortterm_liabilities
return_on_assets range (numeric) afkastningsgrad: result / assets
return_on_equity range (numeric) egenkapitalens forrentning: result / equity
gross_margin range (numeric) dækningsgrad: gross_profit / revenue
operating_margin range (numeric) overskudsgrad: ebit / revenue
has_parsed_financials boolean the company has at least one XBRL-bearing, parsed report
latest_filing_period range (date) period_end of the most recent report

The six ratios use the exact financial_key_figures formulas, each nullif()- guarded on its denominator. gross_margin and operating_margin are null for most companies because turnover is not disclosed.

missing != zero. A range filter needs a non-null value inside the bound. A company with no parsed report is excluded by any figure, ratio, or latest_filing_period filter — never read as zero. A company whose filing does not disclose an input is excluded by a filter on that figure or on a ratio derived from it (a null depreciation makes ebitda null; a null revenue makes both margins null). has_parsed_financials is the one determinate financial key.

Out of scope. Multi-year trend predicates ("positive gross-profit trend over three years") are not Filterable fields — the surface exposes the latest period only, the same scope note #58 recorded. A filing older than the retained five-year window (ADR-0007) has no row at all, so latest_filing_period filters within that window.

Web classification filters

The public contract accepts these source-qualified website observations:

Key Shape Meaning
sales_areas Set: local, regional, national, international Quoted service or delivery coverage with named geography.
target_markets Set: b2b, b2c, b2g Quoted company statements about business, consumer, or public-sector customers.
technologies Set: wordpress, woocommerce, shopify, google_tag_manager, google_tag, meta_pixel Approved installation signatures in retained HTML.
has_webshop Boolean An identified Shopify or WooCommerce product-add form.
has_google_tag Boolean A Google loader and a matching configured tag identifier.
has_facebook_pixel Boolean A Meta loader and a literal Pixel initialization call.

ADR-0025 fixes the source rules and detector catalog. Each set uses OR within the selection. Different keys use AND. Unknown enum values are rejected. Empty sets add no condition. A false boolean means no approved signature was found in the recorded bounded HTML coverage. It does not mean absence across the entire website. Missing, failed, conflicting, expired, and identity-scoped HTML evidence cannot match either boolean selection. Tag Manager alone does not establish a Google tag.

company_detail(cvr).current_detail.web_classifications exposes each key's current evaluation, latest_attempt, and append-only history. Evaluations include eligibility, freshness, value, source URLs and locators, coverage, confidence where applicable, rule/extractor/prompt/model versions, and source fetch and expiry times. The underlying evidence table is not granted to clients.

The frozen page_fresh_seconds policy currently gives evidence 30 days from the earliest source fetch. Re-extraction does not renew this period. A failed refresh preserves earlier eligible evidence and records the failed attempt. A newer successful negative or conflict supersedes earlier positives. Expired evidence remains visible in detail but is excluded from search and exact counts. Publication writes and expiry boundaries invalidate search versions. Search itself does not fetch pages. Existing free-text profile values are not converted into filter values; normal refresh or compatible re-extraction must produce the new versioned outputs.

The consumer implementation is sales.sourceagent#187. On a retained fixture with 2,260,001 companies and 60,006 synthetic website evaluations, industry/municipality queries with Google tag, national coverage, or combined technology/tag selections took 0.021–0.049 seconds for search and 0.475–0.533 seconds for exact facets.

Company contact availability

The Kontakt group uses the existing company flags has_registry_phone, has_registry_email, has_registry_website, and advertising_protected, plus these fields:

Key Shape Meaning
registry_roles Set: board, executive_management Any current qualifying natural-person CVR membership.
has_attributed_work_contact Boolean At least one retained, unsuppressed contact attributed to a company-scoped person mention.
has_board_work_contact Boolean Such a contact with an unreversed identity link to this company's current board member.

The role mapping and privacy boundary are fixed in ADR-0024. Role dates come from each member's FUNKTION, not its organisation container. Both endpoints are inclusive on the UTC read date. Unknown participant kinds, legal entities, and board alternates do not match. Unknown role values are rejected. An empty set adds no condition. Other filter keys combine with AND.

company_detail(cvr).current_detail.contact_availability publishes the same roles and nullable flags, the rule version, registry provenance, and the people evaluation in the latest relevant run with its input coverage and versions. It returns no personal contact values or suppression reason.

True requires retained supporting evidence. False requires a successful latest people evaluation over a nonempty recorded bounded text input and no retained supporting value. Board false also requires an evaluable registry collection. False refers to this retained dataset, not all people or pages in the world. Uncollected, failed, or unevaluable routes yield null without positive evidence; both boolean selections exclude null. Later collection failure does not remove retained positive contact evidence. Contacts have no age-based expiry under ADR-0015. Availability does not mean verified identity, verified delivery, or permission to contact someone.

Suppression, contact deletion, identity-link correction, collection results, and membership changes invalidate search versions in the same transaction. Counts remain exact and never duplicate a company for multiple contacts or roles. Normal CVR collection fills the new membership projection; older rows remain unavailable until refreshed. Search does not acquire new evidence. The existing general role-history date defect is tracked in issue #720.

Employee and decision-maker contact filters are unsupported. A website title cannot establish those relationships. The consumer must use these public keys and must not join personal-data base tables.

The consumer change is sales.sourceagent#185. Its live contract conformance suite must pass before delivery. On the retained 2,260,001-company fixture, an industry and municipality query with attributed contact presence took 0.058 seconds for search and 0.725 seconds for exact facets. The equivalent board-role query took 0.030 and 0.521 seconds. The internal presence predicate reads only the required evidence and reuses its CVR-keyed plans; building full detail provenance per candidate took 5.50 seconds for contact facets. Broader unselective query work remains in #634.

Current registry auditors

auditor_ids is an any-of set of strings from company_auditor_vocabulary(). Each ID is a CVR enhedsNummer for a company participant, not a CVR number or an audit-firm name. The vocabulary supplies the latest disclosed name. An empty set adds no condition; a nonempty set requires a matching current registry firm. Unknown or ended IDs return no matches. Malformed identifiers are rejected. All other filters combine with AND. Counts remain exact and count a company once, including joint auditors.

company_detail(cvr).current_detail.auditors contains current firms, the full dated collection from the current document (history), and retained source publications. Each publication includes source update, collection time, and raw reference. state=unavailable means the relationship collection was not collected; state=collected includes an empty collection. Empty does not claim that no real-world auditor exists. unavailable entries explain unsupported participant kinds, missing identity, or missing membership starts. Natural persons remain source evidence and cannot match the firm filter.

Current means FUNKTION=REVISION membership started on or before the UTC read date and has no end or an end on or after that date. Organisation-container dates do not substitute for membership dates. CVR registry appointments are separate from filing audit firms, signing persons, and submitting enterprises. See ADR-0023.

CVR delta writes and population swaps publish appointments atomically and keep old source snapshots in the restricted auditor-publication table. Source changes and UTC date changes expire search cursors. Normal CVR refresh or reconciliation fills old rows; until then their auditor evidence is unavailable. Search never contacts CVR. No additional source or credentials are required. The sales consumer change is sales.sourceagent#183. Its live filter conformance suite must pass before release.

The opt-in auditor performance fixture contains 2,260,000 synthetic companies and auditor evidence for one fifth of them. The verified industry, municipality, and auditor query returned one company: search took 0.074 seconds and exact facets took 0.515 seconds. Before the query changes, search took 1.134 seconds and scanned 2,034,529 nonmatching active names. Binding the four facet arrays before planning lets PostgreSQL use the industry index. Reusing the internal CVR-keyed plan reduced repeated planning in exact facets. Auditor-only queries remain subject to the broader query work tracked in #634.

Gross-profit growth

gross_profit_growth is a growth object, for example {"years": 3, "min": 20}. Both fields are required. years is 1 or 3; min is a finite percentage with an inclusive lower bound. No other child keys are accepted. company_search_filter_contract().object_schema declares this shape and the allowed horizons. Other filter shapes have a null object_schema.

The percentage is 100 * (latest - baseline) / baseline. A baseline must be positive: 100 to 120 is 20%; 100 to -20 is -120%. A missing, zero, or negative baseline is unavailable. This is endpoint change, not CAGR or a claim that profit grew in each intervening year.

Select the latest retained accounting period, then its latest publication, before looking for metrics. A newer unparsed or PDF-only correction makes that endpoint unavailable. The baseline ends exactly N calendar years earlier; Postgres normalizes a leap day to the last day of February. Both periods must be annual and their start dates must also differ by exactly N calendar years. Both endpoints must fall within the existing five-year retention window. Missing intervening periods do not prevent comparison.

The extractor retains the fact's ISO4217 currency, entity identifier, start date, and non-dimensional scope. Both endpoints must use the same currency and entity and match their filing periods. Unverified or conflicting metadata is unavailable. No currency conversion or consolidation inference occurs. Old extraction rows become eligible only after the existing metrics backfill re-extracts them with version v5. No unit or scope is filled from a default.

The predicate applies to search_companies and every company_facet_counts bucket and exact total. It has no separate bucket. Other keys combine with AND; existing disjunctive facet rules remain unchanged.

company_detail(cvr).current_detail.financials.gross_profit_growth contains entries 1 and 3. Each reports percentage, absence_reason, rule_version, and both endpoint objects with publication ID, dates, value, comparison metadata, source timestamps, parsing time, taxonomy/extractor version, and raw source reference. A missing endpoint is null. Reasons distinguish missing endpoints, unavailable metrics, nonpositive baseline, incompatible periods, currency, scope, fact periods, nonfinite metrics, and the retention cutoff.

Financial mutations advance a derived publication revision in the same transaction. The search version also includes the UTC date to bind the rolling retention cutoff. Previous pages and facet contexts expire after either changes. Bulk replacement changes the report relation identity atomically and preserves the publication triggers for later mutations.

Migration: 20260910120000_gross_profit_growth.sql. Decision: gross-profit endpoint comparisons.

Filterable fields: web presence

Status: implemented (#630, spec #626). Company search can filter on the web-enrichment presence output. All five keys are determinate booleans applied by company_search_matches_cvr_keyed_filters(cvr, filters). has_website also has an exact indexed facet. The other four keys select matching CVRs before the facet operation counts their index memberships.

Key Shape Source
has_website boolean an eligible current website in company_web_presence, backed by an accepted ownership decision and retained primary-page reference
has_linkedin boolean kind linkedin
has_facebook boolean kind facebook
has_instagram boolean kind instagram
has_web_contact boolean any web_company_contact value for the company

There is no missing != zero case: a company either has a resolved signal of that kind or it does not. Enrichment is explicit-shortlist only (CONTEXT.md), so has_*: true selects within the enriched shortlist. has_website: false also includes a company with only unverified historical directory observations.

Out of scope. The observed profile signals (line_of_business, value_proposition, and the rest) stay out of the Filterable field list for the reason the registry section gives: they are confidence-ranked model observations, not predicates.

Bulk participants contract

Status: implemented (#569, parent #428). The surface is the granted view company_participants, readable by authenticated over PostgREST. It answers leadership at scale: one row per current registry role across the whole population, so a consumer fetches the participants of a page of search_companies results with one cvr=in.(...) filter instead of one company_detail call per CVR.

Columns: cvr, person_id, name, kind, relationship_type, role_name, role_function, ownership_pct, voting_pct, valid_from.

ownership_pct and voting_pct are fractions in [0,1], despite their _pct suffix: 1.0 means 100%, 0.25 means 25%, and 0 means a stated zero. Null means no share was stated. Multiply by 100 only for percentage display. These are source values, not values already scaled to a 0–100 percentage.

  • person_id is the registry enhedsNummer of the participant — stable across role changes. kind is person, company, or other, so a company shown as legal owner reads like any other participant.
  • relationship_type is the edge type the registry states: legal-owner-of, beneficial-owner-of, has-voting-rights-in, and holds-role. The qualifiers are the source's own statement about the assertion, not about either entity.
  • Current roles only. valid_to is null is the view's own definition. An ended assertion is history: company_detail's relationship_history carries it with its period, source, and evidence. A past role never appears here.
  • name and kind can be null: a participant removed from the identity projection keeps their edges, so the relationship is current while the identity projection is gone. A null name is not an empty name, and a null kind is not the value other.
  • The view is live. It carries no data_version; every read reflects the data as of that moment. A consumer needing the same identified snapshot as the grid uses the search and Facet contracts, which bind it.
  • Per-role raw-snapshot provenance is not part of this surface. company_detail is the per-company evidence path.

What one company selection returns

Freshness

Every source slice reports its own state under current_detail.freshness: the source, when collection last wrote it, and present or absent. That is what makes an empty section readable — without it a consumer cannot tell "collection has none" from "this slice is unavailable", and would be free to read either as "the company has none".

source_history carries every effective-dated fact version the source stated, each with valid_from, valid_to, is_current, the source update and collection times, and the raw snapshot it was read from. Both parts are always arrays: an empty one means the company has no stored history, which is a different statement from the slice being unavailable. An unknown CVR sets found to false and current_detail to null rather than returning an empty object that reads like a company with no facts.

Stored fact types for a company today: name, status, company_form, main_industry, secondary_industry_1.._3, address, postal_address, phone, fax, email, website, lifecycle, registered_capital, advertising_protected, and employment_annual / _quarterly / _monthly. A production unit stores the same families minus the ones only a legal entity has, plus parent_cvr.

Financials

current_detail.financials states its own window: window_years and period_end_cutoff, then the retained publications newest accounting period first. The cutoff is part of the payload because without it an empty list reads as "this company files nothing" rather than "we do not retain that far back" — and neither means zero.

Every publication carries its accounting period, correction state, document references, promoted metrics, extractor and taxonomy version, and the absence reasons from #105. Superseded publications stay visible with is_current_for_period: false; a correction is a fact about the filing history, not something to filter away.

metrics.ebitda is derived as metrics.ebit + metrics.depreciation, never stored. It is null when either disclosed input is null.

absence_reasons is either a per-metric object or a single string when the whole publication explains itself — "no_xbrl_document" for a PDF-only filing, "not_parsed_yet" when XBRL exists but no metrics row has been written. A publication outside the five-year window has no row at all, so it is simply absent; the stated cutoff is what makes that absence readable.

Web signals

current_detail.web groups the observed values — presence, contacts, and profile_signals — apart from the registry facts beside them. A CVR email is something the company registered; a website email is something we read off a page. The payload shows both and merges neither.

profile_signals is the current best-supported observation per field: line_of_business, products_or_services, value_proposition, target_market, commercial_posture, pricing_evidence, customer_segments, geography. Each carries its source URL, page fetch time, extractor and prompt version, model, confidence, and a bounded evidence excerpt.

Two rules a consumer must not misread:

  • A field with no row was not observed or not extracted. It never means the company lacks it.
  • confidence is the extractor's own self-report, not a calibrated probability. The current view ranks by it — newest extractor version first, confidence only breaking ties within a version — and nothing may present it as a likelihood.

Earlier observations are never deleted: web_company_profile_signal keeps every version's reading of a page so two can be compared against the same evidence.

Entity mentions and discovered relationships

web.entity_mentions holds the companies that retained website or AI Mode evidence names as a customer, supplier, partner, or competitor. A mention is not a relationship: the source gave a name, not a CVR number, and those are different claims.

A mention becomes an observed edge in relationship_history only after a deterministic unique match — one normalized name matching exactly one active company in the hub (#842: an ended company no longer competes). Zero matches or several leave it a mention, with resolution_state: "unresolved". Confidence plays no part in resolution: the model's self-report is about whether it read the page correctly, not about which legal entity a name refers to.

Resolution is reversible. Reversing it ends the edge rather than deleting it, and the mention reads "reversed" rather than reverting to "unresolved" — the assertion was made, and undoing it does not pretend otherwise. The evidence that supported it stays on both.

Each edge also carries its own qualifiers — role_name, role_function, ownership_pct and voting_pct, as the source stated them. Those belong to the assertion, not to either entity, so a participant who is both a director and an owner of the same company has two edges with their own. They were one attributes object until #255; the object repeated its key names on every row and stored an explicit null to say the source stated nothing, which cost 78 of the table's 397 bytes a row for 289 distinct values across 1.2M rows.

Company-to-company ownership comes from Vrvirksomhed.deltagerRelation, not from the participant surface — that index holds natural persons only. A company participant carries no CVR number, so it is resolved through enhedsNummer, a registry key that is unique in companies. An owner the hub has not loaded yet produces no edge; the next reconcile re-reads the same document and resolves it then.

Registry, observed, and derived edges keep their source_class in the response for their whole life. A web-source claim and a register's statement are never presented as the same kind of thing.

Website people

current_detail.web.people holds every person mention found on the company's website. An empty list means that collection has no stored mention for the company. It is never a missing field and never an error.

Each mention has the name, title, and company role as the website stated them. It also has its evidence, confidence, source URL, page fetch time, raw artifact reference, extractor version, prompt version, and model. An attributed work contact point stays under that mention. Each contact has its own evidence, source URL, page fetch time, and extractor provenance.

source_class: "observed" identifies the website claim. resolution_state is unlinked, resolved, or reversed. An unlinked mention stays in the list with equal standing. A resolved mention has a registry_person object with the matched registry identity. A reversed mention stays visible, but its registry_person is null because the link is no longer current.

Registry participants stay in current_detail.participants. Website mentions stay in current_detail.web.people. The operation does not merge the two source classes. It does not expose person_mentions, person_mention_contact, person_contact_suppression, extractor results, or LLM operation records as separate read contracts.

Participants

current_detail.participants holds the people and entities holding a role now, each with their identity, kind, current name, and roles. An ended role is not there; it is an ended edge in relationship_history.

Role assertions use distinct edge types, because the source states distinct claims: legal-owner-of (the REGISTER register), beneficial-owner-of (HVIDVASK), has-voting-rights-in, and holds-role for everything else. Ownership and voting are separate edges — a share class can carry one without the other.

ownership_pct and voting_pct use fractions in [0,1] throughout participant roles and relationship history. For example, 1.0 is 100% and 0.25 is 25%; null is unknown, not zero. The mapper parses the CVR source strings without rescaling. The semantic quality rule warns on an out-of-range share but does not clip, convert, or discard the source assertion.

tests/test_semantic_rules.py::test_ownership_and_voting_are_fractions_between_zero_and_one checks the range and warning behavior. tests/test_cvr_mappers.py::test_maps_fractional_source_share_without_converting_it checks that the source fraction is not rescaled.

flowchart LR
    CVR["CVR share string: 1.0"] --> Mapper["Parse numeric fraction"]
    Mapper --> Ledger["Ownership or voting edge: 1.0"]
    Mapper --> Quality["Warn when outside [0,1]"]
    Ledger --> Reads["Company and person reads: 1.0"]
    Reads --> Display["Consumer percentage display: 100%"]

ADR-0006 bounds this slice: a natural person contributes identity, kind, current name, and their business roles. Their own stored history is the participant kind and nothing else, because historical names are exactly what the ADR excludes. people_roles is gone; company_participants is a view over the relationship ledger, so there is no second copy of the truth.

current_detail.production_units holds the units the company operates now. A unit it operated in the past is not there; it is an ended operates-production-unit edge in relationship_history. That separation is the point of the model: production_units.cvr is a projection of the current edge, not separate relationship truth, so a P-number reassigned to another company moves the column while both companies keep their edge.

Two rules worth stating plainly:

  • A hemmelig contact point is never stored. CVR marks contact values the company asked not to have published. They do not appear in the current projection, in history, or in a retained artifact.
  • Only the newest employment observation per series is stored. The monthly series runs to dozens of entries per company; the full series stays in the raw snapshot. employees_band reads the annual series, which real companies leave empty — latest_employment_* is the one to use.

Graph-neighbor contract

Status: implemented (#572, parent #428). The PostgREST operation is:

company_graph_neighbors(
  p_cvr bigint[],
  p_min_confidence numeric default 50,
  p_cursor jsonb default null
) returns jsonb

Only authenticated can execute it. The function is stable, security definer, and has the fixed search path public, pg_temp. It is a granted read over the relationship ledger — the graph is the stored current observed edges, not a projection, and there is no publisher job.

It answers competitor traversal for the sales product: one or more CVR anchors in, neighboring companies out. The product proposes competitor candidates from this contract and does not infer new competitor edges.

Traversal rules:

  • Types and classes. Only current (valid_to is null) observed edges of competitor-of and customer-of are traversed. An ended assertion is history — company_detail's relationship_history — and registry edges are never neighbors, even if a mapper ever wrote such a type.
  • Symmetric. An edge observed in either direction makes each company the other's neighbor. Each supporting edge keeps its own observation direction: anchor-is-subject reads "the anchor's page named the neighbor"; anchor-is-object reads "the neighbor's page named the anchor".
  • One row per neighbor CVR. With multiple anchors, the neighborhoods union: a company connected to several anchors is one row whose supporting_edges carries every supporting anchor edge. Anchors themselves are never candidates.

Each item contains neighbor_cvr, neighbor_name, competitor_confidence, and supporting_edges. Every supporting edge carries anchor_cvr, direction, relationship_type, confidence (0-100), observed_at, source_url, excerpt, and extractor_version.

Confidence and threshold rules:

  • competitor_confidence is round(max edge confidence x 100) over the current observed edges between the neighbor and the anchors — the highest-support rule the consumer ranks by. No combined or averaged score exists in the contract, and the value stays the extractor's self-report, not a calibrated probability.
  • The threshold selects neighbors, never evidence. The default 50 (a null argument means the default) keeps a neighbor when its best current edge reaches it; supporting_edges then shows every current edge between the pair, including edges below the threshold.
  • Sorted by competitor_confidence desc, then neighbor CVR asc, in fixed 100-row keyset pages. The opaque cursor binds anchors, threshold, and position; a cursor for a different request is an error.

Error rules:

  • An empty anchor array is an invalid-parameter error.
  • An anchor that does not resolve to a company node is an invalid-parameter error. The consumer resolves anchors to legal-company CVRs through search_companies first, so a non-resolving anchor is caller error, not an empty result.
  • A minimum confidence outside 0-100 is an invalid-parameter error.

The operation is live: it carries no data_version, the same posture as company_participants. A page fetched after a web-enrichment run may mix versions, and the cursor never expires from data changes. Snapshot identity remains a search and Facet concern.

What was removed

company_registry_detail is gone. It was introduced in #215 to publish the registry facts the mapper had started promoting, before the operation carried them; it now returned a strict subset of current_detail, and two surfaces for one question is how they drift.

No version-suffixed compatibility view was added in its place. Consumers and collection change together, which is the posture recorded under Change posture.

The population-filter and financial views stay — active_companies, active_companies_financial_snapshot, financial_key_figures — because they answer questions across companies, which a per-CVR operation cannot.

The population-filter and financial views stay — active_companies, active_companies_financial_snapshot, financial_key_figures — because they answer questions across companies, which a per-CVR operation cannot.

Hybrid-similarity contract

Status: implemented (#587, spec #584, ADR-0018). The PostgREST operation is:

hybrid_similar_companies(
  p_seed_cvr bigint[],
  p_filters jsonb default null,
  p_cursor jsonb default null
)
It is security definer, stable, with a pinned search path, granted to authenticated and revoked from public. The sales application calls it when a user opens, refreshes, or pages the Prospect queue (consumer ADR-0005); it sends approved Seed CVRs and its active policy's filters, and applies tenant decisions to the returned page.

Ranking is Reciprocal Rank Fusion (k = 60, algorithm version hybrid-1.1) over two lists:

  • Semantic: cosine distance between the candidate's vector and each seed's vector at the contract's embedding version — the version most companies carry. Each candidate takes its best seed; other seeds never lower a result. A seed without a current vector still anchors the structured list.
  • Structured: industry group (the first two digits of the registry branch code), employee band proximity (ordered employee groups within one), and revenue band proximity (DKK bands within one), counted against the best seed.

Each row carries the candidate CVR and name, the global hybrid_rank, semantic_similarity (1 − best distance, 6 decimals), best_seed_cvr, the structured matches that fired, the profile-signal evidence references, profile_version, and signal_freshness (the newest page fetch behind the candidate's profile).

Hard filters apply before ranking — a filtered-out candidate is invisible, not down-ranked:

Filter Meaning
industry_group_in, industry_group_ex branch-code group allow/deny lists
employees_min, employees_max ordered employee groups, 1–5: zero, 1–9, 10–49, 50–199, 200+; null or unknown source codes are invisible under the filter
revenue_min, revenue_max DKK from the latest filing; a company without a filing is invisible under the filter
exclude_cvr explicit exclusions
municipality_ex JSON integer array, each code 1–9999, matched against the current registry municipality code; missing municipality is not a match
keyword_ex At most 100 strings, each 1–200 characters after trimming; literal substring exclusions over current company name and profile text

Municipality and keyword exclusions

accepted_prospect_evidence_contract() publishes both rules, their evidence attributes, bounds, missing-data rules and cursor_filter_binding=exact_json. The forward migration is 20260927100000_hybrid_policy_exclusions.sql. Municipality is companies.municipality_code, the current CVR registry code, not a postal code, name or inferred location. The numeric domain validates the code representation; it is not a claim that every integer is an assigned code.

Keywords examine companies.name and each current company_profile_signals.value for line_of_business, products_or_services, value_proposition, target_market, commercial_posture, pricing_evidence, customer_segments and geography. Each value is matched separately. URLs, evidence excerpts, historical observations and vectors are not keyword text. Missing values and profile text restricted by the person-evidence policy are not matches. A name can still match when profile text is missing or restricted.

Trim ASCII spaces, tabs, CR and LF from each keyword, normalize both keyword and source text to Unicode NFC, and apply PostgreSQL lower using the database locale. Match the resulting literal substring with strpos. Accents and internal whitespace remain significant. %, _ and regular-expression characters are literal; there is no stemming or word-boundary rule. See the PostgreSQL string functions. A blank keyword, wrong JSON type or bound violation returns 22023.

Any matching keyword or municipality excludes the company. Empty arrays do nothing. All hard filters apply together before both ranking lists, page selection and accepted-card lookup. An excluded company cannot reappear on a later page under the same current evidence. The accepted card reports outside_context with no rank for the same policy. Available cards include name and nullable municipality_code in their registry evidence, with source times; their current profile observations already carry values and evidence. The consumer reads these functions and has no new table grants.

The cursor binds the exact filter JSON. Changing an array's order, letter case or spacing requires a new first page, even when matching would be equivalent. Object key order is immaterial to JSONB equality. The read is live, so changes to source evidence can change results between pages; the cursor is not a snapshot.

An unknown filter key, a reversed range, an empty seed array, an unknown seed CVR, and a request whose seeds carry no usable Semantic company profile at the contract's embedding version are invalid-parameter errors (22023) — the last is the consumer's own market activation gate, and there is no structured-only fallback.

Candidates without a current vector stay outside the ranking. The response counts them in pending_embedding_count, an unfiltered backlog count — companies with profile signals whose embedding is not at the contract's version yet. They join the queue when the pipeline stage or the backfill stores their vector.

The operation is live: no data_version, the same posture as company_graph_neighbors. The keyset cursor binds the canonical seeds, the filters, the algorithm version, and the position; a continuation must repeat the same seeds and filters. Ranks are global, so page two continues the numbering.

Implemented view surface

The following views are the current implementation from issues #26, #58, and

152. They stay accurate until the replacement migration lands.

View Base table(s) Filter baked in
active_companies companies ended_at is null (active only)
financial_reports_latest financial_reports newest publication per (cvr, period_start, period_end), correction-resolved
active_companies_financial_snapshot active_companies, financial_reports, financial_metrics active + at least one parsed report; that company's single most recent accounting period
company_web_presence web_company_presence, companies, website evidence bundles/pages current best eligible signal per (cvr, kind)
company_web_contact web_company_contact, companies every distinct contact value per (cvr, kind)

active_companies columns: cvr, name, main_industry, municipality_code, postal_code, employees_band, started_at, municipality_name, main_industry_text. The industry text is the source Danish label, or null when the source has none.

active_companies_financial_snapshot columns: cvr, name, main_industry, municipality_code, postal_code, employees_band, financial_period_end, gross_profit, revenue, result, equity, ebit, ebitda, assets, profit_before_tax, staff_costs, depreciation, current_assets, shortterm_liabilities, cash, employees. A company absent from this view has no parsed nøgletal at all (no XBRL, outside the backfill window, or not yet delta-fetched) — never assume absence means zero.

ebitda here is derived (ebit + depreciation), not stored — see financial_key_figures below. employees is the exact average headcount the company reported in its filing, present in about half of filings; employees_band beside it is CVR's band and is always present. They are different things and neither substitutes for the other.

financial_key_figures columns: cvr, load_id, period_start, period_end, ebitda, equity_ratio (soliditetsgrad), current_ratio (likviditetsgrad 1), return_on_assets (afkastningsgrad), return_on_equity (egenkapitalens forrentning), capacity_ratio (kapacitetsgrad), gross_margin (dækningsgrad), operating_margin (overskudsgrad). One row per parsed report, not per company — join to financial_reports for the period, or use the snapshot view for "latest per company". Every formula reproduces proff.dk's published figures for the same filing. A null means an input was not disclosed, never zero, so gross_margin and operating_margin are null for most companies because turnover is not disclosed (ADR-0007).

Filing history is bounded to five accounting years (#190, ADR-0007's 2026-08-12 amendment). financial_reports retains only publications whose period_end falls inside the window — the same predicate the nøgletal backfill selects on — so a company's filings from before the cutoff are absent from every view above, not merely unparsed. Absence therefore means "outside the retained window" as well as "not filed digitally"; it never means zero. Run snapshots contain only records selected by the five-year source filter. Fetched XBRL documents in that window stay in GCS. The upstream offentliggoerelser index is credential-free and unbounded, so widening the window is a new metadata and document backfill with an earlier cutoff.

Why a financial metric is missing

A null metric used to conflate six different situations. Each one now has a name, reported per metric by the operator-facing financial_metric_coverage view (#105). Only the last two mean the extractor has a problem.

Reason Meaning
no_xbrl_document The publication is PDF-only (~37% of filings). Nothing was parsed.
not_parsed_yet An XBRL document exists, but no metrics row has been written for it yet.
fact_absent The filing was read, and it does not disclose this fact. Danish ÅRL lets the smallest reporting class omit most of the income statement.
ambiguous_context More than one whole-enterprise value for this concept and period. The extractor refuses to guess; a rule is missing.
parse_failed The XBRL document could not be read at all.
reason_not_recorded The row was parsed before reasons existed. Reasons are not backfilled.

A publication outside the five-year window (ADR-0007) has no financial_reports row at all, so it is absent from the view rather than carrying a reason. None of these ever means zero.

"Region" maps to municipality_code (kommunekode), not postal_code — CVR has no field literally named "region"; municipality is the coarser filter the charter's M1 phrasing implies. Consumers needing finer geography use postal_code, also exposed on the view.

Financial-band filtering uses gross_profit, not revenue — a deliberate deviation from the PRD/charter's literal "revenue band" phrasing.

54's taxonomy spike found that Danish ÅRL lets the smallest reporting class

(Klasse B/micro — most of the corpus) omit turnover entirely from their published statement, so filtering on revenue would silently exclude most companies. gross_profit is the more universally present top-line figure; revenue is still exposed for consumers who specifically want it (nullable, same "missing != zero" rule).

Enforcement

Every base table (companies, production_units, people, ingestion_runs, financial_reports, financial_metrics) has row-level security enabled with no policies, which blocks any non-owner role regardless of table-level grants — the ingester's own connection (table owner) is unaffected. Only the implemented views are granted to the authenticated role. A consumer with an authenticated Supabase session can query a granted view; querying a base table directly returns zero rows regardless of what it's granted, by construction.

This posture is preserved automatically across every companies/ financial_reports backfill and reconcile run (ingesters/cvr/loader.py's bulk_load_rows): the staging+swap those runs use doesn't natively carry a table's views, grants, or RLS state onto the new relation (Postgres binds all three by OID, not name, and LIKE ... INCLUDING ALL doesn't cover RLS or views at all) — the loader captures and re-establishes all three as part of every swap. financial_metrics is a sibling table with an FK to financial_reports.load_id that never itself participates in the swap; Regnskab's reconcile additionally deletes any financial_metrics row for a report the new scroll no longer contains, before the swap runs (pipeline._delete_orphaned_metrics, _run's pre_swap_hook) — otherwise bulk_load_rows re-adding the FK constraint after the swap would fail on the orphaned reference.

Change posture

  • Read contracts are released with their consumers. A replacement migration changes the current contract in place and removes obsolete surfaces. The project does not retain duplicate versioned views for backward compatibility.
  • A contract replacement is not complete until collection migrations, tests, documentation, and known consumers agree on the same shape. The Known consumer register says which conformance suites must go green for a given surface.
  • PostgREST exposes whatever the public schema grants allow — only the views are granted, so only they are queryable via the API regardless of PostgREST's own configuration.
  • Base-table columns and internal joins are not a consumer contract.

Performance (M1 acceptance criterion)

The M1-style query — active companies by industry × region × employee band — is the read contract's validation gate (PROJECT_CHARTER.md's M1 criterion, company-filter half).

select cvr, name from active_companies
where main_industry = %s and municipality_code = %s and employees_band = %s;

Backed by a composite partial index:

create index companies_m1_filter_idx
  on companies (main_industry, municipality_code, employees_band)
  where ended_at is null;

Measured (tests/test_read_contract_performance.py, opt-in via RUN_PERFORMANCE_TESTS=1 — not part of the routine suite, since generating the population takes real time): 2,260,000 synthetic rows (procedurally generated locally — no live upstream call was made or is needed for this), matching companies' ADR-0005 population estimate. Query plan: Index Scan on the composite index, 0.25ms execution time, 1.6ms measured round-trip including Python/network overhead — comfortably under the <1s target.

Performance (financial-band filter, #58)

select cvr, name from active_companies_financial_snapshot
where main_industry = %s and municipality_code = %s
and gross_profit between %s and %s;

Supported by companies_m1_filter_idx (company-side selectivity) plus:

create index financial_metrics_gross_profit_idx
  on financial_metrics (gross_profit)
  where gross_profit is not null;

The view is defined as a lateral join, not a plain distinct on over the whole financial_reports/financial_metrics join — that was the first attempt, and it measured 3.7 seconds at scale, over the <1s target: a distinct on forces Postgres to resolve "most recent period per company" across the entire population before an outer where on industry/ municipality/gross_profit can apply at all. lateral instead runs the per-company "most recent period" lookup only for companies already matching the outer industry/municipality predicate — Postgres inlines the view and pushes that predicate down to companies_m1_filter_idx first, the same way it already does for active_companies alone, so the lateral subquery only ever executes for a small, pre-filtered set of companies. Kept here as a record of what was tried and rejected, not just what shipped.

Measured (tests/test_regnskab_read_contract_performance.py, same opt-in RUN_PERFORMANCE_TESTS=1 pattern): 2,260,000 synthetic companies + 1,000,000 synthetic financial reports/metrics (procedurally generated locally — no live upstream call was made or is needed for this), yielding 900,538 companies with at least one parsed report. Query plan: Bitmap Index Scan on companies_main_industry_municipality_code_employees_idx narrows to the matching companies first, then a Nested Loop runs the lateral lookup via Index Scan on financial_reports_cvr_period_start_period_end_idx and financial_metrics_pkey for just that narrow set — 0.4ms execution time, 3.1ms measured round-trip including Python/network overhead — comfortably under the <1s target, and consistent with the M1-only filter's own measured result above.

Honest scope gap vs. the charter's exact M1 wording: PROJECT_CHARTER.md asks for "positive gross-profit trend over 3 years" — a multi-year comparison. active_companies_financial_snapshot exposes each company's single most recent accounting period only, not a trend across periods; a "positive trend" query would need to compare multiple financial_reports rows per company (or a dedicated trend view) and isn't built here. This issue's own acceptance criteria describe a single "revenue/gross-profit band" filter, which is what's delivered — the trend clause is a residual gap against the charter's original phrasing, not silently claimed as done.

Known cosmetic quirk: Postgres's LIKE ... INCLUDING INDEXES (used by every staging+swap) regenerates index names from their columns rather than preserving the name a migration gave them — companies_m1_filter_idx becomes companies_main_industry_municipality_code_employees_idx after the first swap. The index itself (columns, partial predicate, and the query plan choosing it) is unaffected; only the name changes. Already tolerated by #20's own index-preservation test. Not fixed here: correlating pre- and post-swap index identity by definition rather than name is a real generalization (analogous to the FK/view fixes already in bulk_load_rows) but adds complexity for a cosmetic-only concern — worth revisiting if a future need (e.g. pg_stat_user_indexes monitoring by name) makes the name load-bearing.

Performance (widened filters, #630)

search_companies with filters and no text query walks companies in lower(name), cvr order and stops at the first 101 that pass. The predicate that must reduce the walked set is the registry filter. Layer 1 routed every registry key through one company_search_matches_registry_filters(...) call; that function does not inline (sub-selects in its body), so the four set-shaped keys — company_form_codes, postal_codes, secondary_industries, latest_employment_bands — could not drive the partial indexes created for them, and a selective filtered browse scanned every active company (66 s at 720k rows in testing).

630 applies those four inline in search_companies and

company_facet_counts as col = any(array(select …)) / && clauses. Under the plpgsql custom plan the jsonb_array_length(…) = 0 guard folds away and the index is used; company_search_matches_registry_filters (now the residual range and boolean keys only) and company_search_matches_cvr_keyed_filters (financial and web-presence, cost 1000 so it sorts last in the filter) run only on the index-narrowed rows. The financial keys keep the #58 LATERAL — "latest accounting period per company", executed for the narrowed set, not the population.

Measured (tests/test_company_search_filter_performance.py, opt-in RUN_PERFORMANCE_TESTS=1): 2,260,000 synthetic companies + 1,000,000 synthetic financial reports/metrics (procedurally generated locally — no live upstream call). Filters: company_form_codes + one postal_codes value + gross_profit band + equity_ratio floor + latest_filing_period floor. 110 ms execution time, 170 ms measured round-trip including Python/network overhead — comfortably under the <1 s target and consistent with #58's own result. company_facet_counts on the same request stays under the same bound.

For a request with only a financial range or has_parsed_financials, company_facet_counts computes the latest financial row and the parsed-CVR set once. It then semi-joins those CVRs to the active-company scan. It does not run the latest-period LATERAL once per company. Facet selections apply to the indexed dimensions after this financial constraint. Text requests evaluate the financial predicate on their indexed search candidates. The population test records the filter-only time separately and requires it to stay below 10 seconds for 2,260,000 companies and 1,000,000 reports (#634).

Web enrichment (#152)

These views expose current presence and company-contact observations. Profile signals and entity mentions have separate read surfaces.

company_web_presence columns: cvr, name, kind, url, confidence, evidence, step, extractor_version, retrieved_at. One row per CVR and signal kind (website, linkedin, facebook, instagram) — the current best signal, not the history.

company_web_contact columns: cvr, name, kind, value, source_url, extractor_version, page_fetched_at. One row per CVR, kind, and value: a company legitimately publishes several phone numbers, and each is its own signal. Personal contact values never reach this view — they are dropped inside the extractor under ADR-0006.

Which row wins

A website is eligible only when its presence row matches the evidence bundle's company, run, primary URL, and verified domain. The bundle must record an accepted ownership decision from a supported lead source, must not mark the page as a listing or parked site, and must retain a successful primary-page reference with its SHA-256 digest. An accepted related or unclear judgment remains accepted under the existing ownership policy; this read does not make a new ownership decision.

An old high-confidence llm-verdict row without that proof remains historical evidence, not a current first-party website (#707). Among eligible rows, confidence decides first, retrieval time second, and row ID breaks ties. Non-website signal selection is unchanged.

Company detail, company search, has_website filtering, and website facet counts use this same selector. The migration corrects affected indexed facet memberships and advances the search publication revision. Pre-cutover cursors expire. It does not delete or rewrite observations, bundles, or page artifacts.

tests/test_company_website_selection.py covers accepted versus unverified history, rejected/listing/parked evidence, accepted related sites, and unchanged social signals. test_company_search_contract.py::test_search_and_facets_exclude_unverified_directory_websites covers the search and facet boundaries.

For the two historical records in #707, run uv run python scripts/audit_website_presence.py. By default it reads the test database secret and checks the retained GCS primary objects. It uses a read-only database connection, bounded records and object sizes, and prints no page bodies or contact values. DATABASE_URL can select another authorized audit database. The output reports the actual current selection and retained history, not a simulated policy. Version-1 evidence without a recorded object generation can prove a current byte-hash match, but cannot prove the original generation.

Absence means unresolved, not absent from the web

A company missing from company_web_presence has no eligible current signal. It may have no website, an unreachable site, no collection request, or only unverified historical observations. The outcome and reason of each run remain internal (company_enrichment_run), not part of this contract.

Operation tables stay internal

company_enrichment_run, company_web_request, source_retrieval, and the evidence bundle tables are how the pipeline runs and what it retained. They are execution mechanics, not findings; publishing them would make run logs and provider task identifiers a contract we would then have to keep. Row-level security gives the client roles no rows.

Shape and measured performance

distinct on (cvr, kind) over the whole table, unlike the LATERAL shape

58 needed for financials. CONTEXT.md constrains paid and scraped

enrichment to explicit shortlists, so this reads thousands of rows rather than the 2.3M-row population that made distinct on too slow there.

Measured on a synthetic shortlist history of 269,789 presence rows over 50,000 companies (tests/test_web_read_contract_performance.py, PG14): a single company's lookup plans onto web_company_presence_current_idx (cvr, kind, confidence desc, retrieved_at desc) and executes in 0.068 ms. The index is what keeps the sort local to one company instead of the population.

This historical timing predates the #707 ownership-evidence checks. It is not a measurement of the revised selector.

The revised selector passed the same 50,000-company shortlist test on local PostgreSQL 15, with 269,789 historical presence rows and one ownership-backed website per company. The first measured lookup returned three signal kinds in 0.092 s, within the test's 1 s budget. A subsequent EXPLAIN ANALYZE reported 0.559 ms execution, using the presence, bundle, and page indexes. These are synthetic local measurements, not live-service latency guarantees.

Exact website evidence references

Website contacts, profile signals, company mentions, person mentions, and person contacts expose evidence_locator through their existing read surfaces and company_detail. It contains retrieval_id, artifact (URI, generation, digest, media type), projection_version, source_url, kind, start, and end. Ranges use Unicode character offsets, with an exclusive end, in the retained page-text projection. A supporting-page claim cites that supporting page. The primary page is not a default citation. Earlier observations without exact references have a null locator; they are not eligible sealed extraction inputs.

These remain source-qualified observations. A search citation is not an acquired website page, and an unresolved company mention is not a verified relationship.

Search source signals

company_source_signals is readable by authenticated consumers. Each row is an extractor's observation of a retained organic or AI Mode source, with cvr, goal, value, self-reported confidence, exact evidence, source_kind, retrieval_id, acquired_at, retrieved_at and extractor target. evidence_locator identifies the immutable raw artifact, JSON pointer and text range. source_url is the provider's result URL when present; it is empty when not supplied. No URL is invented from a query identity.

The surface keeps observations from each extraction run. An absent goal means no signal was published; it does not mean the company lacks that capability or relationship. Read run_id and request progress to distinguish failed and empty extractions. These claims are separate from page contacts, verified website presence, registry facts and canonical company relationships. AI citations do not prove that the cited page was acquired. The input retains full answers, blocks, citations and provider metadata. Current contact objections are applied again when a parsed source result is reused.

Company search dimensions

The following additional keys work in search_companies and in the exact company_facet_counts.total under the same data_version (#615):

Key Input Meaning
production_unit_count Integer range, for example {"min":2} Distinct P-numbers with an active membership in the latest collected CVR penheder array.
purpose_text Nonblank string, up to 500 characters Case-insensitive literal substring of the current CVR FORMÅL text. % and _ are literal characters.
address_changed_since ISO date, for example "2025-09-01" A change between two registered location addresses, effective on or after the date and no later than today.
has_email Boolean At least one collected web company email contact.
has_phone Boolean At least one collected web company phone contact.

Different keys use AND. has_email and has_phone follow the existing has_web_contact convention: false means no retained matching observation. They do not inspect registry contacts or prove that no contact exists online. Use has_registry_email and has_registry_phone for registry evidence.

The first known address is not an address change. Repeated identical address periods do not count; an apartment or floor change does. Dates use UTC and are inclusive. Future changes are excluded. address_changed_since rejects invalid dates and relative date words.

A collected empty production unit list means zero. Missing source data and unreadable current membership dates remain unknown and do not match a range. Expired and future memberships do not count. A conflicting current purpose has no searchable value. New source fields are populated by a normal CVR refresh; this migration does not fabricate values for existing rows.

Production unit count, purpose text, and address changes are filter-only. Email and phone expose exact boolean facet buckets. Company form, website, and advertising protection also expose per-value counts. Each bucket removes only its own filter. All other filters, the text query, and the result data_version still apply.

Company structure and financial classes

ADR-0027 defines these collection-owned filters and their evidence rules:

Key Input Public meaning
is_holding true, false, or "unknown" Registered holding activity under the source's DB07 or DB25 industry contract.
group_role Array of parent, subsidiary, independent, unknown Observed majority voting control. independent means “No registered majority-control link” and requires complete coverage.
performance_class Array of high_growth, liquid, profitable_cash_rich, financially_challenged High gross-profit growth; Current assets cover short-term liabilities; Profitable with cash coverage; Negative equity or a loss with a current-asset shortfall.

Values within a set use OR. Different keys use AND. Parent and subsidiary can both match one company. Financial classes can also overlap. All three keys have exact disjunctive facets. Holding and group facets retain an explicit unknown bucket. Financial facets count true matches only; unknown financial inputs never become a positive match.

Collection retains industry versions, voting intervals and source locators. An unbounded successful CVR population run plus complete company relation collections establishes group coverage. Missing organisation data, unresolved votes, contradictory controlling parents, and control cycles prevent an independence claim. Natural-person ownership does not create a company parent.

Financial extraction version 6 retains monetary facts with entity, context, currency, period, document hash and source locator. Incompatible or ambiguous facts remain unknown. The newest report revision takes precedence even when unreadable. Financial windows require consecutive annual periods and expire 18 calendar months after the latest eligible period. No currency conversion or estimate fills a gap. Re-extraction of the same document keeps its original fetch time.

company_detail.current_detail.classifications returns the rule version, evaluation date, publication version, truth values, reasons, and source evidence. Holding and control use precomputed effective intervals. Financial classes use a stored projection with validity dates. Reads select the UTC date; they do not traverse control graphs or rebuild financial histories. Source writes and full population swaps refresh these projections in the same transaction. The result version also binds the evaluation day.

Existing rows need a normal CVR refresh and financial re-extraction to acquire these source inputs. Until then, unavailable evidence remains unknown. The full-population performance proof passes the 500 ms p95 limit with populated control and financial evidence (#770). All six sales consumer conformance tests pass against this migration chain (consumer PR #189, issue #190). Consumer CI still reads producer main and must pass again after the producer contract is merged. These checks do not constitute a deployment.

Test hub anonymous sign-in

On 2026-09-12 the test hub's /auth/v1/settings response reported anonymous_enabled: true and disable_signup: false (#671). Anonymous sign-in was already enabled, so this work did not change the setting. This check did not create a user or test a browser sign-in.

An anonymous sign-in creates an authenticated session. It is distinct from an unauthenticated request with the public API key. Existing RLS and grants still govern reads; see the Supabase anonymous sign-in documentation.

Company-form vocabulary

company_form_vocabulary() is granted to authenticated and returns code and nullable name, ordered lexically by code. It enumerates distinct non-null company-form codes from retained CVR company records, including ended companies. Names are the source-disclosed Danish short descriptions already collected with those codes. The latest source-updated nonblank name wins; the lowest CVR breaks an equal timestamp. If no record discloses a name, name is null.

This is the collected vocabulary, not a complete legal code catalog. The service does not invent labels or infer that an absent code is legally retired. A code remains listed while retained company records use it. After it disappears from the dataset, a restored selection must retain its code and show its missing name. company_form_codes continues to accept that selection; no current matches means an empty result. Consumers read no base tables and keep no copied code catalog.

The vocabulary supplies names, not facet counts. company_facet_counts supplies exact context-dependent counts under the search publication version. Missing or pending counts must not be rendered as zero. The company-form vocabulary itself is a current read and is not bound to a search cursor.

Person list and profile

Issues #768 and consumer #116/#161; ADR-0028. person_read_contract() publishes the live filter definitions, response field sets, page size/order, query bound, and identity/name/role/contact/claim states. Search validates keys, set sizes, and role values from this operation. A consumer conformance test must compare its handlers with this live vocabulary.

Both data operations are stable, security-definer reads with a pinned search path. Only authenticated has execution access. Private tables and helper views have no consumer grant.

select search_people(
  p_query := 'hansen',
  p_filters := '{"company_cvrs":[41527080],"relationship_types":["holds-role"]}',
  p_cursor := null
);
select person_profile(p_person_id := 4000000001);

The list contains CVR PERSON participants only. person_id is the registry enhedsNummer, unchanged across names, roles, companies, and publications. Company participants and unknown kinds are excluded. Website mentions remain company-scoped observations, not list identities. Unresolved mentions remain available through company_detail; they are never silently merged here.

Search accepts a trimmed, case-insensitive literal substring of the current registry name, or an exact decimal person ID. % and _ are literal text. Null or blank means no text restriction. The limit is 200 characters. It does not search old names, website names, contact values, or titles.

Filter Meaning
company_cvrs 1–100 eight-digit numeric CVRs; a current registry role at one of these companies.
relationship_types 1–100 values from holds-role, legal-owner-of, beneficial-owner-of, has-voting-rights-in.

Values combine with OR; keys combine with AND on the same role. Duplicate values and set order do not change the request. Unknown keys, invalid types, empty sets, malformed cursors, and invalid values raise SQLSTATE 22023. No filters means all known registry person identities, including those whose roles have ended. These filters do not claim employment or decision authority.

The response has items, total_count, data_version, and next_cursor. It sorts by ascending numeric person_id, with 100 items per page. Each item has person_id, name, name_state, state, current_roles, and work_contact_state. A blank or absent name is null with name_state=unavailable. The contact state is available, withdrawn, or unavailable; list rows carry no contact values. Multiple roles never duplicate a person row.

The ownership and voting shares in current_roles and profile role_history use the same fractional unit as Participants: ownership_pct = 1.0 or voting_pct = 1.0 means 100%. Null means no stated share. Neither the person reads nor the mapper multiplies these values by 100.

Pass the opaque next_cursor unchanged with the same query and filters. It binds the last person ID and a transactional publication version. Person, role, identity, company, contact, suppression, and mention changes invalidate it, as does the UTC date. Changes to other company publications may also invalidate it. An expired cursor raises 22023 with person cursor expired; restart at the first page. Totals and rows use one database statement snapshot. Uncommitted changes are not visible. There is no historical snapshot service.

person_profile returns person_id, state, identity, data_version, current_roles, role_history, website_claim_state, website_mentions, and work_contacts. An unknown or non-person ID returns state=unavailable, null identity, and empty evidence collections. If the current identity row was removed but retained participant-kind evidence identifies a person, the ID remains readable with state=withdrawn and no current name. History is retained.

The identity contains the current registry name, kind, source, collection time, and field evidence for the name. The current-name projection does not retain a source update timestamp: source_updated_at is explicitly null. This read adds neither historical names nor residential data. participant_facts_as_of already takes a person ID and date; it remains the narrow kind-history read.

Each role carries its company CVR, current company name or unavailable state, type, source role/function, ownership/voting percentages, inclusive source period, and source provenance. evidence_id is a stable hash of the source assertion's natural key, not a confidence score. Correcting its values or end date does not change that key. Different source records remain separate. state is current, ended, future, or unknown_period. A missing start cannot prove current membership. role_history includes every retained registry assertion; current_roles contains only those current on the UTC date. Every field of an assertion shares its source record and collection evidence.

Website mentions expose their own name/title/role and per-field evidence, source URL, fetched/recorded times, extractor/prompt/model versions, and self-reported confidence. They never replace registry fields. Distinct claims within one company remain separate and give website_claim_state=conflicting. Different roles at different companies are not a conflict. A reversed link returns only mention_id and state=withdrawn, not the old person's claim.

work_contacts.items contains attributed work email/phone values permitted by ADR-0008/0015, each with contact ID, mention ID, CVR, source evidence and state. A suppression excludes a value before deletion. Deletion retains only contact and mention IDs; the published item becomes {contact_id, mention_id, cvr, state: "withdrawn"}. The aggregate is available if any permitted value remains, withdrawn if only withdrawn references remain, otherwise unavailable. Missing means not available in retained evidence, not proof of no real-world contact. Deleting a whole mention also removes its contact references. Pre-migration contact deletions cannot be reconstructed and remain unavailable.

Free-text person evidence and URLs may repeat contact values. Once a company has a suppression or retained withdrawal, those unstructured evidence fields are withheld for that company's person claims and contacts. Typed permitted values and source times remain readable. No suppression value, reason, note, raw personal artifact, or contact-entitlement state is returned.

Accepted-prospect evidence

Issue #769 and consumer #163; ADR-0028.

select accepted_prospect_evidence(
  p_cvr := 41527080,
  p_seed_cvr := array[12345678]::bigint[],
  p_filters := '{"industry_group_in":["62"]}',
  p_expected_versions := null,
  p_linked_person_id := null
);

accepted_prospect_evidence_contract() publishes the current ranking versions, accepted filter rules, seed bounds, response fields, missing-reason vocabulary, and card/linked-person states. Queue and card validation share these definitions; consumers can check their handlers against them without reading private tables.

This authenticated read directly selects one CVR after the global Hybrid calculation. It does not enumerate or page through the runtime queue. Use the original target-market seeds and hard filters. Do not pass the queue's current seller-state suppression list as that context: an explicitly excluded CVR remains outside the context. The tenant store keeps references and seller state; it does not copy platform evidence.

The result is evidence_mode=current, with read_at, cvr, state, reason_code, context, current_versions, ranking, reason, and linked_person. It does not reproduce the rank at acceptance. The ranking and reason can change when source evidence or the candidate population changes. The optional expected versions contain exactly algorithm and embedding from a prior Hybrid response. They are policy expectations, not a snapshot ID.

State Meaning
available The CVR currently ranks in this saved context.
unavailable reason_code identifies company_unavailable, seed_unavailable, seed_profile_unavailable, profile_unavailable, or outside_context. Rank/reason are null.
withdrawn A supplied prior version reference can no longer be evaluated because a required company, seed, or profile is missing. No prior evidence is replayed.
policy_changed ranking_version_changed: supplied algorithm/embedding versions differ from the current versions. Rank/reason are null; the caller can explicitly request the current policy.

The rank is the ordinal position in the saved context, never a percentage or a sale probability. It shares hybrid-1.1, RRF with k=60, and the dominant embedding version with hybrid_similar_companies. Hard filters apply before ranking; the direct lookup applies after ranking. The response names the best seed, semantic similarity, versions, and rank meaning. Queue and card reads select tied financial periods by period start, publication date, then load ID, so corrections and their evidence have a deterministic order.

The reason has code hybrid_similarity, a factual description of the ranking basis, structured_matches, and supporting evidence: registry company facts with source times, selected financial publications and revenue, stored embedding versions and creation times, and current profile field observations with URLs and fetch times. A stored vector's profile_version can differ from the latest profile observations; the separate versions/times make this visible. A missing optional fact is not zero or a negative finding. This reason is a similarity basis, not a new model of willingness to buy or permission to contact.

A linked person is never inferred from registry participants. A caller may supply a seller-selected p_linked_person_id. The read returns that registry reference only if an unreversed mention for this CVR has a currently permitted attributed contact. The object has state, person_id, and mention_ids, with no contact values. No selection or unsupported evidence is unavailable. Reversed or withdrawn supporting contact evidence gives withdrawn with a null person ID. Several supported mentions can qualify the same selected person. This does not choose a person, buy a contact, or grant an entitlement.

Seeds must contain 1–100 eight-digit CVRs. Filter keys match the existing Hybrid contract; malformed sets/ranges and unknown keys raise 22023 before data availability is checked. These validation rules also apply to the queue read. The private shared ranking routine, person helpers, and base tables have no consumer grant. The authenticated producer conformance suites cover direct lookup beyond 100 candidates, shared ranks, source evidence, missing states, changed policy, selected-person scope, and withdrawal.

Profile evidence values and URLs in accepted-card evidence are withheld after a company contact suppression or retained withdrawal, because these fields can repeat a removed contact. Their state is unavailable; source times and versions remain readable. The queue also withholds affected profile evidence URLs. The same private privacy predicate governs these reads and person reads.