Imported from zhenzonglin/WMH (
AGENTS.md). Install upstream withnpx skills add zhenzonglin/WMH. Copyright stays with the author.
WMH analysis project
-
Current default work is recovery only.
wmh-study startruns only the historical recovery report;wmh-study recovery-path --through analyseruns the separate all-stroke extension without Word or HTML. BP, CEC, and kidney remain archived and run only when explicitly selected. The prior four-study contract, code, run directories, and result pointers remain intact; do not combine extension results with the historical four-test Holm family. Keep quantitative WMH continuous and publish no clinical cutoffs. Seedocs/recovery_path_workstation.md. -
Current work is under
src/wmh_hcy/studies,docs/four_studies_plan.md, andwmh-study. H_CHD missing values follow user rules H_HD=1 -> 0 and observed H_CHD_TP -> 1. Preserve recorded values and provenance; contradictory evidence blocks BP fitting and must never be imputed away. Auxiliaries do not enter the regression. CEC retains HDL-C adjustment; M0 shares the primary completed data. Kidney keeps all three albuminuria interactions in both multinomial equations with a single primary persistent-category interaction. Usedocs/four_studies_workstation.md; the recurrence/Hcy notes below are historical. -
Historical recurrence workflow is
wmh-hcy recurrenceunderdocs/recurrence_v3_plan.md(2026-09-17, after reviewing year-one results). One Y5_IS/Y5_IS_DD cohort, seven truncation horizons, shared core MI and frozen design; no death regression, absolute risk or bootstrap. Preserve legacy H1-H4/longterm files and pointers. Its output root is recurrence_v3. See recurrence_v3_workstation.md for its startup/audit instructions; the remaining bullets apply to legacy workflows unless stated otherwise. -
Work in the existing clone. All statistical code is Python 3.11. Support either
conda env create -f environment.ymloruv sync --frozen. Conda uses the version pins exported fromuv.lockintorequirements-conda.txt; do not independently upgrade analysis dependencies. -
Patient data are external and read-only. Never commit raw SAS, patient CSV, images, ID maps, local configurations, or outputs. Public SAS fixtures under
tests/fixtures/sasare the only bundled SAS data. -
The workstation configuration is
config/workstation.local.yml. Start withwmh-hcy audit; accept the actual SuStaIn path withconfigure --sustain-dir; then run throughpreparebefore fitting models. -
Follow H1 through H4 and
docs/statistical_analysis_plan.md. Use only the source fields infields.py; do not invent unavailable variables, infer a center, change the endpoint, silently drop required covariates, or substitute synthetic data for real input. -
Review
outputs/real/audit/andprepared/flow.csvfor data availability. Per the researcher's 2026-09-16 revision, manual image review labels do not determine eligibility in any analysis; do not reintroduce a QC gate or relabel records as passed. Preserve exact IDs and reasons for data exclusions. -
Keep method definitions and CLI documentation consistent. In the activated Conda environment run
ruff check src testsandpython -m pytest -q; for uv prefix those commands withuv run. Add targeted verification for changed statistical or data-handling behavior. -
Report software validation separately from patient-level results. Synthetic effects and counts are not research findings.
-
The researcher added cumulative 2-5 year IS/IS_DD and mRS endpoints on 2026-09-17. Use the isolated
longtermcommand anddocs/longterm_analysis_plan.md; preserve one-year result pointers. Long-term death dates are unavailable: do not extrapolate one-year censoring or infer death dates from mRS=6.
