Python review performance measurements¶
Issue #785, change 11 (#796). Measured on 2026-09-16 with Python 3.13.15
and the locked development environment. The baseline is 6dc1907.
The revised extractor indexes required facts once and computes currency units
once. Metrics, absence reasons, classification inputs, and source positions
were identical for all three fixtures.
| Fixture | Before, median ms | After, median ms | Ratio |
|---|---|---|---|
aarsrapport_2025.xml |
1.422 | 0.491 | 2.89 |
aarsrapport_2025_inline.xhtml |
0.944 | 0.320 | 2.95 |
lasso_x_2025_real.xml |
10.336 | 3.069 | 3.37 |
Each median uses seven samples of 100 extractions after one warm-up call.
All calls use the fixture bytes and period end 2025-12-31. The measurement
excludes file I/O, database calls, and HTTP. It is a local CPU comparison,
not a production throughput result. The retained root tree and a small fact
index remain in memory; this is not a streaming XML parser.
To repeat, run this code in each checkout with the same Python environment. Compare the complete returned dictionaries between revisions as well as time.
import datetime
import statistics
import time
from pathlib import Path
from da_business_data_collection.ingesters.regnskab.xbrl_extractor import (
extract_financial_metrics,
)
for name in (
"aarsrapport_2025.xml",
"aarsrapport_2025_inline.xhtml",
"lasso_x_2025_real.xml",
):
content = (Path("tests/fixtures/xbrl") / name).read_bytes()
period_end = datetime.date(2025, 12, 31)
result = extract_financial_metrics(content, period_end)
samples = []
for _ in range(7):
start = time.perf_counter()
for _ in range(100):
extract_financial_metrics(content, period_end)
samples.append((time.perf_counter() - start) / 100)
print(name, statistics.median(samples), result)
Change 15 (#799) also reuses fixed Pydantic adapters. It has no measured performance claim; existing request and result-validation tests check its behavior.
Raw storage memory¶
Change 12 (#797) writes JSON fragments into gzip through a temporary spool. The spool moves to disk above 1 MiB. The reader downloads compressed bytes into the same bounded spool, then yields one decoded record at a time. Callers that stop early must close the iterator. Content hashes still cover exact uncompressed UTF-8 NDJSON bytes, with LF between records and no final LF. Uploads still require generation zero.
Measured with tracemalloc on the same Python version and baseline as above:
| Input | Before, peak bytes | After, peak bytes |
|---|---|---|
| 1,000 generated records, each with 16 KiB of text | 33,117,637 | 604,120 |
| One generated record with 8 MiB of text | 27,265,254 | 25,510,237 |
These inputs use repeated x text. The bucket fake discards uploaded data;
its file upload reads 64 KiB at a time. Tracing includes input generation,
serialization, and compression. It excludes GCS transport and measures Python
allocations, not resident process memory. To repeat, wrap each call to
RawStore.write_ndjson_gz with tracemalloc.start() and
tracemalloc.get_traced_memory()[1], using a generator of
{"id": i, "text": "x" * width} for the sizes above.
Batch memory no longer scales with the full uncompressed extract. A large single string still requires source, encoded JSON, and UTF-8 allocations; the large-page result shows this limit. The reader also needs enough memory for one complete decoded record. Temporary disk space must hold the compressed object. This change does not claim bounded memory independent of record size.
Deferred ingestion memory¶
Change 14 (#798) keeps rejected documents in a temporary JSON spool with a 1 MiB memory threshold. It sends at most 128 rejection rows per database call, after COPY has ended. The run still owns commit and rollback.
Ownership assertions must wait until the company swap releases its lock. A private temporary SQLite table retains the latest document per source ID. Its index supplies sorted IDs without a Python list of the whole population. Publication reads at most 2,000 documents per chunk, as before. Empty corrections, native dates, and decimal percentages are preserved. The run owner closes queued work on failure; publication closes the temporary database on success and failure. Internal pickle values are read only from this private spool, never from downloaded artifacts.
This uses SQLite's temporary database behavior with disk spill enabled and a 1 MiB pager cache target. Database work still follows Psycopg's COPY transaction rules. The design does not query a connection while COPY owns it.
| 1,000 deferred records, each with 16 KiB of generated text | Before, peak bytes | After, peak bytes |
|---|---|---|
| Rejected raw records | 16,774,751 | 1,109,073 |
| Ownership documents | 16,942,805 | 69,117 |
Measurements use tracemalloc around record generation and intake. Each reject
has a name field. Each ownership document has a distinct ID, one assertion
with a date and Decimal percentage, and the generated text in its artifact URI.
The database fake receives no calls during intake. These are synthetic Python
allocation measurements, not production RSS or SQLite native allocations.
A regression test also checks 4,001 assertion documents across publication chunks.
Spool capacity still scales with the population. A temporary filesystem backed by RAM counts against the container memory limit; spilling does not remove that cost. One unusually large record can also exceed the normal batch memory size.