Skip to content

Architecture

The collection service is the platform boundary for upstream access and the only writer to the Supabase data hub. Its units of operation are independent single-source ingesters and composite enrichment pipelines; both persist checkpointed outputs through the store rather than calling each other directly.

The repository-local decision spine starts with ADR-0011. When a user selects a company, the collection-owned detail contract returns current detail, source history, and relationship history as defined by ADR-0012. Relationships use the relational graph in ADR-0013, and web-derived facts follow ADR-0014.

The web enrichment architecture runs one company pass per Cloud Tasks task in one service: search, AI Mode, and the company website in, evidenced signals out. ADR-0029 records that decision.

See PROJECT_CHARTER.md at the repository root for the source registry and delivery scope.

Shared registry execution (#782)

registry.pipeline owns bounded scroll execution, checkpoints, quality gates, and publication. registry.loader owns generic upserts and staged swaps. CVR and Regnskab own their contracts, mapping rules, and entity-specific hooks. Neither ingester imports the other. Runtime entry points sit above orchestration; the import checks enforce these layers.

Typed financial extraction

XBRL contexts, promoted numeric columns, absence codes, and classification evidence have static shapes. Metric values remain Decimal or null. Evidence retains its source positions, periods, currency, and digest. The pipeline keeps these types through parsing and adds provenance before the SQL boundary. At that boundary, classification evidence uses Jsonb; an empty absence map remains SQL NULL.

Registry static contracts

EntityConfig[RecordT] binds the record contract, row mapper, fact mapper, relationship mapper, and assertion mapper to the same validated source model. Shared execution keeps that type parameter. The company default is explicit in its parameter type; other source models remain generic. Database callbacks use psycopg.Connection, document period ends use datetime.date, and report IDs use strings.

DocFetcher.fetch_and_parse[T] returns the parse hook's result type. The scroll protocol requires both a closeable generator and retry counters. Controlled clients state zero HTTP attempts; their transport gate is therefore skipped. A static regression test checks valid inference and rejects both a mismatched entity mapper and a wrong document-parser return type before execution.

Typed request and extraction models

Company-request normalization keeps the existing Pydantic request, selection, and target models until the admission serialization boundary. Normalization copies an input model before it sorts or deduplicates values. The stored canonical JSON retains the same optional-field shape used for idempotency checks. After-model validators return Self. Extraction output and generated schemas use a recursive JSON value contract; no second runtime validation hierarchy is added.

Typed decision traces

Trace writes require a read-only identity protocol with a nullable run ID. Backfills can therefore state that they have no run. Run-bound operations use the stricter derived protocol with an integer run ID. Static trace row shapes continue through the SQL insert, while decoded detail remains recursive JSON. The existing detail length limit and truncation format are unchanged.

Typed history records

History mappers return static fact, relationship, and ownership assertion shapes. The history writer retains these types through buffering and publication. Exact tuple types describe the fact and relationship unique keys. Deferred ownership assertions wait in the unlogged table assertion_staging as JSON; the writer restores native dates and Decimal percentages when it reads them back. Source membership periods expose fixed qualifier types without adding runtime models.

Typed company receipts

Company admission and inspection return static receipt shapes. Nested progress records retain native SQL timestamps, numeric IDs, explicit nulls, and decimal cost strings. The stored JSON policy and evidence remain JSON boundaries. Trusted SQL reads are typed at the selected row shape; no second runtime model hierarchy is introduced. Work handlers share one static response contract. Routes keep the existing generic dictionary response encoder so static typing does not remove fields or change status codes, idempotency, or post-admission dispatch.

Typed registry projections

CVR mappers return fixed company, production-unit, person, and role row types. They retain integer IDs, nullable parent CVRs, native dates, and Decimal capital and percentages. The shared MappedRow contract carries the source timestamp and optional snapshot URI. Financial publication metadata also states its row shape so it can use this same interface. Generic SQL and rule readers accept read-only mappings. Configured column names remain a dynamic lookup boundary; no source document is copied or revalidated to erase its type.

Source annotation coverage

Ruff enforces missing parameter and return annotations under src, including batch_job.py. Test helpers and standalone scripts are outside this rule. Dynamic JSON and SQL boundaries can still use Any; ANN401 is not a blanket prohibition. Database operations use Psycopg types, provider callbacks use their transport contracts, and JavaScript parsing uses Tree-sitter Node and a recursive static syntax type. Repair callbacks have typed clocks, dispatchers, SQL parameters, phase names, and result counters. Nullable SQL reads are checked before access.