Architecture¶
The collection service is the platform boundary for upstream access and the only writer to the Supabase data hub. Its units of operation are independent single-source ingesters and composite enrichment pipelines; both persist checkpointed outputs through the store rather than calling each other directly.
The repository-local decision spine starts with ADR-0011. When a user selects a company, the collection-owned detail contract returns current detail, source history, and relationship history as defined by ADR-0012. Relationships use the relational graph in ADR-0013, and web-derived facts follow ADR-0014.
The web enrichment architecture runs one company pass per Cloud Tasks task in one service: search, AI Mode, and the company website in, evidenced signals out. ADR-0029 records that decision.
See PROJECT_CHARTER.md at the repository root for the source registry and
delivery scope.
Shared registry execution (#782)¶
registry.pipeline owns bounded scroll execution, checkpoints, quality gates,
and publication. registry.loader owns generic upserts and staged swaps.
CVR and Regnskab own their contracts, mapping rules, and entity-specific hooks.
Neither ingester imports the other. Runtime entry points sit above orchestration;
the import checks enforce these layers.
Typed financial extraction¶
XBRL contexts, promoted numeric columns, absence codes, and classification evidence
have static shapes. Metric values remain Decimal or null. Evidence retains its
source positions, periods, currency, and digest. The pipeline keeps these types
through parsing and adds provenance before the SQL boundary. At that boundary,
classification evidence uses Jsonb; an empty absence map remains SQL NULL.
Registry static contracts¶
EntityConfig[RecordT] binds the record contract, row mapper, fact mapper,
relationship mapper, and assertion mapper to the same validated source model.
Shared execution keeps that type parameter. The company default is explicit in
its parameter type; other source models remain generic. Database callbacks use
psycopg.Connection, document period ends use datetime.date, and report IDs
use strings.
DocFetcher.fetch_and_parse[T] returns the parse hook's result type. The scroll
protocol requires both a closeable generator and retry counters. Controlled
clients state zero HTTP attempts; their transport gate is therefore skipped.
A static regression test checks valid inference and rejects both a mismatched
entity mapper and a wrong document-parser return type before execution.
Typed request and extraction models¶
Company-request normalization keeps the existing Pydantic request, selection,
and target models until the admission serialization boundary. Normalization
copies an input model before it sorts or deduplicates values. The stored canonical
JSON retains the same optional-field shape used for idempotency checks.
After-model validators return Self. Extraction output and generated schemas use
a recursive JSON value contract; no second runtime validation hierarchy is added.
Typed decision traces¶
Trace writes require a read-only identity protocol with a nullable run ID. Backfills can therefore state that they have no run. Run-bound operations use the stricter derived protocol with an integer run ID. Static trace row shapes continue through the SQL insert, while decoded detail remains recursive JSON. The existing detail length limit and truncation format are unchanged.
Typed history records¶
History mappers return static fact, relationship, and ownership assertion shapes.
The history writer retains these types through buffering and publication. Exact
tuple types describe the fact and relationship unique keys. Deferred ownership
assertions wait in the unlogged table assertion_staging as JSON; the writer
restores native dates and Decimal percentages when it reads them back. Source
membership periods expose
fixed qualifier types without adding runtime models.
Typed company receipts¶
Company admission and inspection return static receipt shapes. Nested progress records retain native SQL timestamps, numeric IDs, explicit nulls, and decimal cost strings. The stored JSON policy and evidence remain JSON boundaries. Trusted SQL reads are typed at the selected row shape; no second runtime model hierarchy is introduced. Work handlers share one static response contract. Routes keep the existing generic dictionary response encoder so static typing does not remove fields or change status codes, idempotency, or post-admission dispatch.
Typed registry projections¶
CVR mappers return fixed company, production-unit, person, and role row types.
They retain integer IDs, nullable parent CVRs, native dates, and Decimal capital
and percentages. The shared MappedRow contract carries the source timestamp
and optional snapshot URI. Financial publication metadata also states its row
shape so it can use this same interface. Generic SQL and rule readers accept
read-only mappings. Configured column names remain a dynamic lookup boundary;
no source document is copied or revalidated to erase its type.
Source annotation coverage¶
Ruff enforces missing parameter and return annotations under src, including
batch_job.py. Test helpers and standalone scripts are outside this rule. Dynamic
JSON and SQL boundaries can still use Any; ANN401 is not a blanket prohibition.
Database operations use Psycopg types, provider callbacks use their transport
contracts, and JavaScript parsing uses Tree-sitter Node and a recursive static
syntax type. Repair callbacks have typed clocks, dispatchers, SQL parameters,
phase names, and result counters. Nullable SQL reads are checked before access.