Skip to content

Python review performance measurements

Issue #785, change 11 (#796). Measured on 2026-09-16 with Python 3.13.15 and the locked development environment. The baseline is 6dc1907. The revised extractor indexes required facts once and computes currency units once. Metrics, absence reasons, classification inputs, and source positions were identical for all three fixtures.

Fixture Before, median ms After, median ms Ratio
aarsrapport_2025.xml 1.422 0.491 2.89
aarsrapport_2025_inline.xhtml 0.944 0.320 2.95
lasso_x_2025_real.xml 10.336 3.069 3.37

Each median uses seven samples of 100 extractions after one warm-up call. All calls use the fixture bytes and period end 2025-12-31. The measurement excludes file I/O, database calls, and HTTP. It is a local CPU comparison, not a production throughput result. The retained root tree and a small fact index remain in memory; this is not a streaming XML parser.

To repeat, run this code in each checkout with the same Python environment. Compare the complete returned dictionaries between revisions as well as time.

import datetime
import statistics
import time
from pathlib import Path

from da_business_data_collection.ingesters.regnskab.xbrl_extractor import (
    extract_financial_metrics,
)

for name in (
    "aarsrapport_2025.xml",
    "aarsrapport_2025_inline.xhtml",
    "lasso_x_2025_real.xml",
):
    content = (Path("tests/fixtures/xbrl") / name).read_bytes()
    period_end = datetime.date(2025, 12, 31)
    result = extract_financial_metrics(content, period_end)
    samples = []
    for _ in range(7):
        start = time.perf_counter()
        for _ in range(100):
            extract_financial_metrics(content, period_end)
        samples.append((time.perf_counter() - start) / 100)
    print(name, statistics.median(samples), result)

Change 15 (#799) also reuses fixed Pydantic adapters. It has no measured performance claim; existing request and result-validation tests check its behavior.

Raw storage memory

Change 12 (#797) writes JSON fragments into gzip through a temporary spool. The spool moves to disk above 1 MiB. The reader downloads compressed bytes into the same bounded spool, then yields one decoded record at a time. Callers that stop early must close the iterator. Content hashes still cover exact uncompressed UTF-8 NDJSON bytes, with LF between records and no final LF. Uploads still require generation zero.

Measured with tracemalloc on the same Python version and baseline as above:

Input Before, peak bytes After, peak bytes
1,000 generated records, each with 16 KiB of text 33,117,637 604,120
One generated record with 8 MiB of text 27,265,254 25,510,237

These inputs use repeated x text. The bucket fake discards uploaded data; its file upload reads 64 KiB at a time. Tracing includes input generation, serialization, and compression. It excludes GCS transport and measures Python allocations, not resident process memory. To repeat, wrap each call to RawStore.write_ndjson_gz with tracemalloc.start() and tracemalloc.get_traced_memory()[1], using a generator of {"id": i, "text": "x" * width} for the sizes above.

Batch memory no longer scales with the full uncompressed extract. A large single string still requires source, encoded JSON, and UTF-8 allocations; the large-page result shows this limit. The reader also needs enough memory for one complete decoded record. Temporary disk space must hold the compressed object. This change does not claim bounded memory independent of record size.

Deferred ingestion memory

Change 14 (#798) keeps rejected documents in a temporary JSON spool with a 1 MiB memory threshold. It sends at most 128 rejection rows per database call, after COPY has ended. The run still owns commit and rollback.

Ownership assertions must wait until the company swap releases its lock. A private temporary SQLite table retains the latest document per source ID. Its index supplies sorted IDs without a Python list of the whole population. Publication reads at most 2,000 documents per chunk, as before. Empty corrections, native dates, and decimal percentages are preserved. The run owner closes queued work on failure; publication closes the temporary database on success and failure. Internal pickle values are read only from this private spool, never from downloaded artifacts.

This uses SQLite's temporary database behavior with disk spill enabled and a 1 MiB pager cache target. Database work still follows Psycopg's COPY transaction rules. The design does not query a connection while COPY owns it.

1,000 deferred records, each with 16 KiB of generated text Before, peak bytes After, peak bytes
Rejected raw records 16,774,751 1,109,073
Ownership documents 16,942,805 69,117

Measurements use tracemalloc around record generation and intake. Each reject has a name field. Each ownership document has a distinct ID, one assertion with a date and Decimal percentage, and the generated text in its artifact URI. The database fake receives no calls during intake. These are synthetic Python allocation measurements, not production RSS or SQLite native allocations. A regression test also checks 4,001 assertion documents across publication chunks.

Spool capacity still scales with the population. A temporary filesystem backed by RAM counts against the container memory limit; spilling does not remove that cost. One unusually large record can also exceed the normal batch memory size.