Prompt file imported from ZachBach/labOptimal (
.claude/commands/add-analyte.md). Copyright stays with the author.
/add-analyte — widen the panel safely
services/engine/src/laboptimal_engine/data/reference_ranges.py is the source of
truth. LabParser builds its alias index from this table, so an analyte that is
not here can never be extracted — there are 19 today, which notably excludes
Glucose.
Adding one is usually a single entry. The care is in the aliases.
The alias-collision hazard — check this first
LabParser._match_analyte looks for any alias as a substring anywhere in the
line, longest alias first. A short alias therefore matches text that has
nothing to do with the analyte.
This has already caused a live bug: mg is an alias for magnesium, and mg is a
substring of mg/dL. So the first line of any report containing mg/dL — in
practice Glucose or Creatinine — is claimed as magnesium, and the seen dedupe
then means the real Magnesium row is never read. See TASKS.md § First-run
findings.
Before adding any alias, run these checks:
cd services/engine
.venv/Scripts/python.exe -c "
import sys; sys.path.insert(0,'src')
from laboptimal_engine.data.reference_ranges import build_alias_index
NEW = ['your', 'proposed', 'aliases']
idx = build_alias_index()
UNITS = ['mg/dL','ng/mL','pg/mL','g/dL','mmol/L','umol/L','nmol/L','pmol/L',
'mIU/L','IU/L','U/L','ug/dL','mcg/dL','fL','%','x10E3/uL','mg/L','ratio']
for a in NEW:
hits = [u for u in UNITS if a.lower() in u.lower()]
if hits: print(f'COLLISION {a!r} is a substring of units: {hits}')
for existing in idx:
if a.lower() != existing and (a.lower() in existing or existing in a.lower()):
print(f'OVERLAP {a!r} vs existing alias {existing!r} -> {idx[existing]}')
if not hits: print(f'ok {a!r}')
"
Rules that follow from this:
- Avoid aliases shorter than 4 characters unless they are unambiguous
(
hba1cgood,a1cborderline,mgactively harmful). - An alias must not be a substring of any unit string.
- An alias must not be a substring of another analyte's alias, or the longest-first ordering silently decides which one wins.
- Prefer the full printed label a lab actually uses —
"vitamin d, 25-hydroxy"beats"vit d".
Qualifier discipline
Calcium, Ionized is a different test from Calcium, with a different
reference range. The benchmark's truth resolver already refuses to substring-
match across distinguishing qualifiers (ionized, free, direct, fasting,
urine, ratio, …) precisely because the parser makes this mistake.
If the new analyte is a qualified variant, give it its own canonical entry, not an alias on the base analyte.
The entry
ReferenceRange(
canonical="…", # stable key used across the whole system
display_name="…", # human-readable label the UI shows
unit="…", # canonical unit the normalizer converts to
reference_low=…, # standard interval
reference_high=…,
optimal_low=…, # functional band — a curatorial judgement, not a standard
optimal_high=…,
aliases=("…",), # what the parser may see. See the collision check above
nutrients=("…",), # what a shortfall maps to, feeds the recommender
source="…", # REQUIRED — flows into Protocol.citations
),
The source is not optional. Every range must cite where its interval comes
from; the pipeline aggregates these into Protocol.citations so provenance
travels with the result. README § Scope commits to ranges being cited from
published sources rather than established here. An uncited range breaks that
claim. Track sourcing in docs/reference-ranges.md.
If the analyte needs a unit conversion, add it to _CONVERSIONS in
normalize/normalizer.py — only where the conversion is safe and unambiguous.
After the change
cd services/engine
pytest -q # golden file will diff if parsing changed
Expect the golden file to move only if you intended it to. test_golden.py
pins the entire protocol for a fixed panel. If adding an analyte changed the
output for the existing sample panel, that is an alias collision until proven
otherwise — investigate before regenerating. The regeneration command is in that
file's docstring; run it only once the diff is understood.
Then measure the effect on the benchmark:
laboptimal-benchmark run --ocr text
Adding an analyte raises panel coverage and moves the gated denominators — previously out-of-panel truth rows become in-panel and scoreable. Accuracy can legitimately fall as a result, because the system is now being asked to do more. Report coverage and accuracy together; a coverage gain with an accuracy drop is not automatically a regression, but it is never a silent one either.
Checklist
[ ] Alias collision check run — no unit substrings, no overlap with existing aliases
[ ] Qualified variants given their own canonical entry, not an alias
[ ] Reference interval cited; source recorded in docs/reference-ranges.md
[ ] Optimal band justified, or left None
[ ] nutrients mapping set (or empty tuple if none applies)
[ ] Unit conversion added to _CONVERSIONS if needed
[ ] pytest green; any golden-file diff understood before regenerating
[ ] Benchmark re-run; coverage and accuracy both reported
