Operator Runbook¶
Reference for all CLI commands used in this project.
Setup¶
Skills and subagents ship as a repository plugin from seed1207/tooling-skills-agents. Install it before running any skill:
/plugin marketplace add seed1207/tooling-skills-agents
/plugin install tooling-skills-agents@tooling-skills-agents
For a new repository created from this template, then run /setup-project in your
agent. It is the one-time bootstrap flow and must finish before the first
engineering flow. It writes or verifies the issue-tracker, triage-label, and
domain-doc configuration under docs/agents/ that the engineering skills read.
Manual tooling commands used by that flow:
uv sync # Install all Python dependencies
cp .env.example .env # Configure environment variables
bash scripts/link-hooks.sh # Wire git hooks
If the project keeps GCP/OpenTofu delivery, run /bootstrap-gcp-tofu-github
after /setup-project. It provisions the GCP projects and OpenTofu state
bucket, creates Cloudflare Pages with Access, configures GitHub Actions WIF, and
writes the repository secrets needed by CI/CD.
Local variables required before workflow bootstrap:
OPEN_ROUTER_API_KEY=...
CLOUDFLARE_ACCOUNT_ID=...
CLOUDFLARE_API_TOKEN=...
CLOUDFLARE_ALLOWED_EMAIL=you@example.com # Optional; defaults to gh user email
Workflow bootstrap should leave these GitHub repository secrets in place:
WIF_PROVIDER
WIF_SERVICE_ACCOUNT
TF_STATE_BUCKET
GCP_TEST_PROJECT
GCP_PROD_PROJECT
OPEN_ROUTER_API_KEY
CLOUDFLARE_ACCOUNT_ID
CLOUDFLARE_API_TOKEN
Verify the full setup:
bash scripts/doctor.sh
gh workflow run docs.yml
gh workflow run cd.yml -f environment=test -f action=plan
Engineering Flow¶
For a non-trivial feature, first run /grill-with-docs to record settled
language and decisions. If a question needs runnable evidence, use /handoff
to a separate prototype directory, run /prototype, then hand the answer back.
For multi-session work, run /to-spec, then /to-tickets; each ticket records
its blockers. Use /implement per ticket in a fresh context. It drives /tdd
and closes with /code-review. /research is standalone cited reading,
/triage is for incoming requests only, /diagnosing-bugs establishes a
reproduction before fixing, and /wayfinder resolves large decision work.
Quality Gates¶
# Python
uv run ruff check . # Lint
uv run ruff format . # Format (in-place)
uv run ruff format --check . # Format check only
uv run ty check src/ # Type check
uv run pytest # Run all tests
uv run pytest -k "name" # Run specific test
uv run pytest --cov=src # Tests with coverage report
Documentation¶
uv run mkdocs serve # Local preview at http://localhost:8000
uv run mkdocs build # Build static site to site/
Docs deploy exclusively through the GitHub Actions Docs workflow after a
merge to main, or by manual dispatch from Actions -> Docs -> Run workflow.
Do not deploy docs from a local shell.
Issue Tracking (GitHub Issues)¶
gh issue list --label ready-for-agent # List available work
gh issue view <number> # View issue details
gh issue edit <number> --add-label in-progress # Claim an issue
gh issue close <number> # Mark issue complete
gh issue create --title "..." --body "..." # File a new issue
Infrastructure Provisioning (OpenTofu)¶
gh workflow run cd.yml -f environment=test -f action=plan
gh workflow run cd.yml -f environment=test -f action=apply
gh workflow run cd.yml -f environment=prod -f action=plan
Preview and apply infrastructure through the GitHub Actions CD workflow.
Merging to main applies test. Production rollout is explicit through
promote-release.yml.
Do not run local tofu apply, tofu destroy, or docker push for normal
delivery. See infra/README.md for first-time workspace setup and required
variables.
Releases¶
uv run ruff check .
uv run ruff format --check .
uv run ty check src/
uv run pytest
uv run mkdocs build --strict
git tag vX.Y.Z
git push origin vX.Y.Z
gh workflow run promote-release.yml -f version=vX.Y.Z -f dry-run=true
gh workflow run promote-release.yml -f version=vX.Y.Z
Before creating the release, update CHANGELOG.md, bump the package version in
pyproject.toml and src/data_sourceagent/__init__.py, update the
version test, merge that PR from a green main, then create the tag. A tag is
vX.Y.Z or a SemVer pre-release such as vX.Y.Z-alpha; the package version
uses the PEP 440 form of the same version (X.Y.Za0). The tag
creates a prerelease candidate through release.yml; production deploys only
when a maintainer promotes that candidate through promote-release.yml.
Backfill Order¶
The backfill is a sequence, not four independent jobs. Run it in this order.
An entity started before the one it depends on fails on a foreign key, loads
nothing, and records the run failed (#222).
1. cvr / virksomhed companies no dependency
2. cvr / produktionsenhed production_units needs 1
3. regnskab / offentliggoerelser financial_reports needs 1
4. regnskab / metrics-backfill financial_metrics needs 3
cvr / deltager people no dependency, run any time
Only two columns create the constraint, both nullable, both pointing at
companies:
| Child | Column | Set from |
|---|---|---|
production_units |
cvr |
the unit's current virksomhedsrelation |
financial_reports |
cvr |
the publication's cvrNummer |
The order is the whole answer. produktionsenhed and offentliggoerelser
cannot land against an empty hub, only after step 1. Once a companies run has
succeeded, every production-unit and Regnskab run, backfill or delta, reads
only documents updated before the companies watermark, so a unit or
publication for a company registered after that run waits for the next run
instead of failing its foreign key (ADR-0031). The daily schedules keep the
same order: companies and people at 03:00, production units and Regnskab at
04:00; the dependent reconciles run on the second of the month. Both were tried out of
order on 2026-08-15 and both failed on their constraint (#222).
deltager has no foreign key into companies. Participants are their own
population, and their role edges resolve through entity_nodes, which is
created on demand. It may run at any point in the sequence.
Company-to-company ownership edges (#230) resolve inside step 1 and need no separate pass: a full backfill stages the whole population, swaps it, and only then flushes its queued assertions, so every owner is present by the time they are resolved. This is not true of a delta, which flushes per batch — an owner loaded later in the same run produces no edge until the next reconcile.
Watch item: the constraint assumes every referenced CVR exists in the
virksomhed population. That population is unfiltered (ADR-0005 decision 4), so
it should. If step 3 still hits financial_reports_cvr_fkey after step 1
completes, then publications reference companies outside the CVR index, and the
mapper must null those rather than pass them through.
Exercising the Pipeline in Test¶
Restrict a run with a CVR bound: --cvr-from A --cvr-to B. It selects
every document in a block of CVR numbers, the same block for every entity, so
the subset is self-consistent, the foreign keys resolve, the load still stages
and swaps, and the swap guard rails still evaluate.
There is no document cap. --max-docs existed until 2026-08-16 and is deleted.
It took whatever the scroll returned first, so the entities it collected were
unrelated and their keys failed, and it switched the loader to upsert-merge —
a second load path that never swapped and never let a guard rail run. It
answered "does this run at all"; a bound answers "does the pipeline work",
which is the question a test environment exists to settle.
# One CVR block, end to end, in the dependency order above.
BOUND="--cvr-from=10000000 --cvr-to=10999999"
gcloud run jobs execute cvr-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill,$BOUND
gcloud run jobs execute cvr-ingester-production-units --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill,--entity,production-units,$BOUND
gcloud run jobs execute cvr-ingester-people --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill,--entity,people,$BOUND
gcloud run jobs execute regnskab-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.regnskab,backfill,$BOUND
That block held 180,328 documents when it was counted on 2026-08-15 — 34,732 companies, 42,662 production units, 85,973 participants, and 16,961 publications — about 18 minutes end to end.
The bound applies to backfill, delta, and reconcile on both ingesters, and
both ends are required: one end alone would silently mean "to infinity". It
narrows an entity's declared population rather than replacing it, so a
bounded offentliggoerelser run still honours the five-year accounting window.
The Regnskabsdata metrics-backfill command accepts the same bound. It selects
only already-landed reports for CVRs in that range. A bounded financial delta
passes its range to the same metrics selector.
A bounded run drops rows whose current company is outside the bound and
reports them as rows_outside_bound. A production unit's virksomhedsrelation
is historised, so the scroll matches on any parent it ever had while
production_units.cvr holds the current one — 0.4% of a live bounded sample.
They are not rejects; they are simply not in the subset.
A bounded deltager run also drops edges inside a document it keeps, reported
as edges_outside_bound (#254). A participant belongs to the subset while some
of the companies they serve do not, and the scroll filter cannot reach inside a
document to say so. Before this existed, 84.7% of a bounded run's edges named
companies outside the bound and pulled them into entity_nodes with no
projection row — which is what filled the test database (#253).
A bounded run still holds a few company nodes outside its bound, and this is
expected (#262). They are the former operators of production units that are
themselves inside the bound: virksomhedsrelation is historised, so a unit that
is kept still produces one edge per past parent. A bounded run over
10000000-10999999 left 298 such nodes against 6,152 in-bound companies, all
reached only through ended edges, with no facts and no projection row.
They are kept deliberately. "This company used to operate this unit" is real history about an entity that is in scope, and dropping it would erase the only record that the earlier operator existed — the failure ADR-0013 and #219 exist to prevent. The edge is closed, so nothing presents it as current.
So #254's expectation reads: no node outside the bound with an open edge or a fact, not "no node outside the bound".
Size the bound to the environment. Test is a 500 MB Supabase database. The
10000000-10999999 block cost 986 MB on 2026-08-15 and stopped the environment.
Measured cost after #254 is roughly 1.0 KB per document; check
pg_database_size before choosing a block.
Each run records what it covered in ingestion_runs.population_bound
(cvr 10000000-10999999, or null when unbounded), because a document count
means nothing without the population it came from.
The schedule carries the bound too, not only operator runs. ingestion_cvr_bound
in infra/<env>.tfvars appends --cvr-from/--cvr-to to every scheduled delta and
reconcile, for both ingesters. Test sets 10000000-10099999; prod leaves it unset
and covers the whole register (#317).
Set it for every entity or for none. A bounded company load beside an unbounded production-unit load yields units whose parent company is absent, which is the FK failure "Backfill Order" above exists to prevent. One variable feeds all eight scheduled jobs so the entities cannot disagree.
Changing the bound in a populated environment is not a resize. A delta only adds;
it never removes what an earlier, wider bound loaded. Narrowing takes a reconcile
to take effect, and widening loads the new population at backfill cost — check
pg_database_size first.
CVR Ingester Operations¶
Scheduled jobs (infra/batch_job.tf, infra/reconcile_jobs.tf,
infra/staleness_check_job.tf) run the daily deltas, monthly
reconciliation, and hourly freshness check automatically — when
ingestion_schedulers_paused is false. Monthly reconciliation also requires
reconcile_schedulers_paused to be false. All runs use the population
ingestion_cvr_bound names. The two pause controls have safe defaults:
the freshness check asserts a contract only ingestion can satisfy, so it
pauses and resumes with the ingesters. Manual
invocation — e.g. the one-time full backfill, which is deliberately not a
schedule default (ADR-0005) — uses gcloud run jobs execute, overriding
the job's default args:
# One-time full backfill (staging + atomic swap), per entity.
# Order matters — see "Backfill Order" above. Wait for each to succeed.
gcloud run jobs execute cvr-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill
# Only after the companies run has succeeded:
gcloud run jobs execute cvr-ingester-production-units --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill,--entity,production-units
# No dependency; may run at any point:
gcloud run jobs execute cvr-ingester-people --region=<region> \
--args=-m,data_sourceagent.ingesters.cvr,backfill,--entity,people
# On-demand freshness check (same command the hourly schedule runs):
gcloud run jobs execute cvr-ingester-staleness-check --region=<region>
A new environment must have its secrets before its jobs (#138):
OpenTofu creates every runtime identity (cvr-ingester-<env>@<project>,
web-enrichment-<env>@<project> and web-starter-<env>@<project>), every
secret container, and every grant (infra/access.tf, ADR-0030). Only the
values come from infra/set-secrets.sh <project-id>: the registry secrets
(cvr-es-user, cvr-es-password, cvr-database-url) and the five provider
credentials — webshare-proxy-url, openrouter-api-key,
openai-api-key, dataforseo-login, and dataforseo-password. A new
environment therefore takes two applies (see infra/README.md). OpenTofu
refuses to build a job or service before each mounted secret has a version
and its identity has the grant. This is not a theoretical guard: the CVR jobs
were created one day before cvr-database-url existed, every revision came up
Ready=False / SecretsAccessCheckFailed, and because a later apply changed
nothing about the job spec, no revision replaced the broken one for three
weeks. If you ever do find a job stuck that way, changing anything in its spec
forces a new revision and clears it.
Interrupted runs (#139): the same hourly job first closes out any
ingestion_runs row still running after 24 hours. A container killed by
the kernel — an out-of-memory event, for instance — runs no Python
afterwards, so nothing else ever moves that row out of running, and the
ledger silently under-reports failure. The bound is wider than the longest
job timeout, so a legitimately long backfill is never swept. A swept row
carries no watermark, so the next delta still re-reads from the last real
success. The count comes back as interrupted_runs_swept in the job's JSON
output.
Operational health (#25, widened #58, tightened #141): select * from
ingestion_staleness; in Supabase answers "how fresh is each entity?"
directly — across every source (CVR and Regnskab) in one query. A run counts
towards freshness only if it carries a watermark. A process that exits 0
without ever establishing a position upstream is not evidence of anything,
which is how a broken delta query reported virksomhed fresh against an
empty table for three weeks (#156). A quiet-day delta that reads zero
documents still carries the previous watermark forward, so it does count. A failed job execution or a
staleness breach both alert via Cloud Monitoring: by email to every address
in alert_emails (prod: the maintainer, #929), and to Pub/Sub, which
gcp-incidents.yml can mirror into a GitHub issue labeled status,
gcp-incident, <job-slug> when it is run — see infra/monitoring.tf for the alert policies and
docs/adr/0005-cvr-ingestion-single-surface-incremental.md's 2026-07-19
operational-health amendment for the design.
Rebuild registry role history after #720¶
After the migration and service image have passed test CD, execute
cvr-ingester-reconcile-companies, then cvr-ingester-reconcile-people.
Wait for each execution to succeed. Use the deployed job arguments so the
rebuild keeps the environment's CVR bound. Do not enable monthly schedules
to perform this repair.
Bounded rebuilds replace only the declared CVR range. Outside company, production-unit, and financial-report rows retain their original values and cache timestamps. Safety checks evaluate the selected range, and row counts exclude preserved rows. A bounded participant rebuild does not delete global identities absent from its query; only a full-register query can establish that absence. It still replaces each returned document's roles inside the range.
The participant reconciliation job needs the same 2 GiB memory limit as the daily participant job. The measured peak for this dense CVR range was 1465 MB (#258); a 1 GiB reconciliation was killed during the #720 rebuild. Keep the limit in infrastructure code and deploy it through CD before retrying.
The companies job and the companies reconciliation run with 2 CPU and 8 GiB.
A full companies scroll used to spool its ownership assertions to a SQLite file
until the swap, and Cloud Run's filesystem is memory: the first prod backfill
was killed at 1 GiB after about 710,000 of 2.26M companies (#925). The writer
now stages deferred assertions in the unlogged table assertion_staging, 50
records at a time, and deletes its rows when the flush ends or the writer
closes (#926). Container memory no longer grows with the population. Lower the
sizes only after a full prod run shows the peak memory of the job. Every job timeout stays below the 24-hour interrupted-run age,
so the hourly check never sweeps a run that is still working; the reconciles
allow 20 hours.
Each accepted source document replaces its complete registry role history in
one transaction. This removes stale open container-based assertions and
corrected dates. Ended member terms remain available in company_detail.
Company documents rebuild company-to-company assertions; participant documents
rebuild company_participants and the candidates used for person resolution.
An empty assertion set removes earlier assertions from that same document.
Rejected documents retain their previous projection and require separate review.
Verify both job executions and their ingestion_run_report rows. Inspect
ingestion_run_quality warnings and rejects. Compare member periods in retained
source documents with company_detail history, and confirm that ended members
are absent from company_participants. Record execution IDs, row counts, and
any unrepaired rejected documents on #720 before closing it.
Regnskabsdata Ingester Operations¶
Scheduled daily delta and monthly reconcile (infra/regnskab_job.tf) run
automatically. Manual invocation — the one-time full backfill (~6.4M
publications, deliberately not a schedule default per ADR-0007/#55,
mirroring the CVR companies split) — uses gcloud run jobs execute:
# One-time full backfill (staging + atomic swap).
# Requires the CVR companies backfill to have succeeded first — every
# publication carries a `cvr` that must already exist (see "Backfill Order").
gcloud run jobs execute regnskab-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.regnskab,backfill
# Bounded historical metrics backfill — last 5 filing years by default,
# resumable (re-running skips already-parsed reports automatically, #57):
gcloud run jobs execute regnskab-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.regnskab,metrics-backfill
# Widen the window later with a re-run, not a code change:
gcloud run jobs execute regnskab-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.regnskab,metrics-backfill,--cutoff-years,10
financial_reports_latest and active_companies_financial_snapshot in
Supabase are the read contract — see docs/reference/read-contract.md for
the full surface, enforcement, and performance details.
financial_metric_coverage is the operator's per-metric coverage view: one
row per publication and promoted metric, each null carrying a typed reason
(#105). The daily delta also fetches+parses XBRL inline for every publication
with an xbrl_url and upserts financial_metrics (#56) — a publication
without one (PDF-only, ~37% of filings) has no financial_metrics row and
reports no_xbrl_document; absence means "not filed digitally", never zero.
A document that cannot be parsed at all no longer fails the run: it lands a
row whose every metric reads parse_failed, and the raw document stays in
GCS for a later extractor version to re-read.
Monthly reconcile repairs drift/deletion for financial_reports the same
way companies/production-units already do (#58); manual invocation mirrors
the CVR pattern:
gcloud run jobs execute regnskab-ingester --region=<region> \
--args=-m,data_sourceagent.ingesters.regnskab,reconcile
Git Hooks¶
Hooks run automatically: - pre-commit (via prek): ruff lint + format check + ty type check - pre-push: pytest
Data hub sessions (hub-session)¶
Consumers get their Data hub sessions from the Edge Function hub-session
(ADR-0032). _deliver.yml deploys it with every delivery: to test from CD, and
to prod from the promotion. It reads these from the GitHub environment
(test or production), set up under #918:
- secret
SUPABASE_ACCESS_TOKENand variableSUPABASE_PROJECT_REF; - Edge Function secrets in each Data hub project:
TENANT_SUPABASE_URL,TENANT_SUPABASE_PUBLISHABLE_KEY, andALLOWED_ORIGINS(the sales app origins, comma-separated). Supabase providesSUPABASE_URL,SUPABASE_SERVICE_ROLE_KEY, andSUPABASE_ANON_KEY.
Check a deploy: POST https://<ref>.supabase.co/functions/v1/hub-session
without a token must answer 401. The live check is
supabase/functions/hub-session/live_check.ts.
Local checks: in supabase/functions/hub-session, run
deno fmt --check --line-width 100 . && deno lint . && deno check *.ts && deno test --allow-net=none.
Web enrichment¶
Company requests, inspection, re-extraction, and repair are in
Web-enrichment operations. One company
run is one company pass; its outcome, reason, costs, and every step are in
company_enrichment_run.log. Retained raw pages and search responses stay
in the data bucket under raw/web/; nothing deletes them on a timer.