Anton Dziatkovskii · ORCID 0000-0001-7408-3054
Preprint, 10 September 2026. Version of record: Zenodo, DOI 10.5281/zenodo.22696755. License CC BY 4.0.
Full text (PDF, 8 pages, 203 KB)DOICode: blood-panel-pipeline
System paper. Not an evaluation.
A person who has had blood drawn for ten years owns a decade of evidence in a form nothing can query. Almost everyone in this position writes the same parser once, badly, and never publishes it, so the same defects are rediscovered privately every time. This paper describes a small stdlib-only system that turns such an archive into a provenance-tracked fact database, and argues for the decision it is built around: deterministic, testable rules decide what is mechanically true, and a language model, if added later, may only explain. Four planes separate concerns — an immutable verbatim raw layer, a canonical parsed twin rebuilt on every run, a rules plane, and a model plane deliberately shipped empty — and ambiguity is recorded as a quality-control flag on the row rather than resolved silently or dropped.
The empirical content is a failure, reproducible from the published artifact alone with no private data: mapping markers to standard LOINC identifiers by substring assigns a wrong code to 12 of 52 matched analytes, including three distinct measurements carrying the code for haemoglobin and red blood cells in urine carrying the code for red blood cells in blood. The instructive part is the repair that did not work — re-ordering candidate keys longest-first turned the failing test green while changing the outcome on exactly one row of 158, so a green suite certified a class that was 92% open. The real repair was to abandon substring matching: exact normalised match only, no code at all for an unmatched analyte, assigned codes dropping from 52 to 40.
Two principles generalise: for medical identifiers missing is recoverable and wrong is not, and a fix that turns a red test green is evidence about an instance, never about a class. The paper separates what a reader can run from what is the author's word, names the studies that would turn this case report into a measurement, and reports two defects found in the author's own artifact while writing it.
Privacy: the repository ships no measurements and neither does this paper. No measured value, reference range from a real draw, draw date or diagnosis appears anywhere in it.
Reproducibility: every figure about the system is re-derivable by running the public repository; tests/reproduce_miscoding.py re-derives the miscoding table from the two shipped mapping files and exits non-zero if the published figures stop reproducing.
This is not medical advice and the system described is not a diagnostic device.
Code: github.com/tonydzi/blood-panel-pipeline (MIT) · Author: github.com/tonydzi · tonydzi.github.io · ORCID 0000-0001-7408-3054 · dzyatkovskiy.a2@gmail.com
personal health informatics · provenance · LOINC · UCUM · clinical data normalization · privacy by design · self-tracking · deterministic rules · large language models · reproducibility · laboratory data · SQLite
Dziatkovskii, A. (2026). Facts Before Explanation: A Provenance-Tracked Pipeline for a Decade of Personal Lab Panels. Preprint. Zenodo. https://doi.org/10.5281/zenodo.22696755
@misc{dziatkovskii2026bloodpanelpi,
author = {Dziatkovskii, Anton},
title = {Facts Before Explanation: A Provenance-Tracked Pipeline for a Decade of Personal Lab Panels},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22696755},
url = {https://doi.org/10.5281/zenodo.22696755},
note = {Preprint, CC BY 4.0}
}
One text, one DOI: the PDF on this page is the Zenodo file byte for byte. Please cite the DOI.