Imported from wolski/diann-runner (
AGENTS.md). Install upstream withnpx skills add wolski/diann-runner. Copyright stays with the author.
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Overview
This is a Python package for running DIA-NN (mass spectrometry data analysis) workflows. It provides:
- A Docker wrapper (
diann-docker) to run DIA-NN in containers - A Snakemake workflow for automated pipeline execution
- QC plotting utilities
The package uses uv for dependency management and cyclopts for CLI argument parsing.
Git Workflow
- Work directly on
master; do not create local feature branches unless the user explicitly asks for one. - Before changing files, check the current branch and working tree status.
CLI Entry Points
All commands defined in pyproject.toml:
| Command | Module | Purpose |
|---|---|---|
diann-docker |
diann_docker.py |
Run DIA-NN in Docker container |
run-diann |
run_diann_cli.py |
Normalize AppRunner/SUSHI inputs → run the workflow (apprunner/sushi subcommands) |
diann-snakemake |
snakemake_cli.py |
Run Snakemake workflow (low-level passthrough) |
diann-cleanup |
cleanup.py |
Alias for snakemake --delete-all-output |
diann-qc |
plotter.py |
Generate QC plots from DIA-NN results |
diann-qc-report |
qc_report.py |
Generate Markdown QC report |
thermoraw |
thermoraw_docker.py |
Convert .raw files via Docker |
prolfquapp-docker |
prolfquapp_docker.py |
Run prolfqua QC in Docker |
prozor-diann |
prozor_diann.py |
Re-run protein inference on DIA-NN report |
Optional (in contrib/oktoberfest/): Koina/Oktoberfest integration for alternative spectral predictors (not included in main workflow).
Development Commands
Installation
cd diann_runner
uv venv
source .venv/bin/activate # or `.venv/Scripts/activate` on Windows
uv pip install -e .
# For testing:
uv pip install -e ".[test]"
Testing
Unit Tests:
# Run tests directly
python3 tests/test_workflow.py
# Or with pytest
python3 -m pytest tests/
Snakemake Workflow Execution:
Use diann-snakemake to execute the workflow:
# Navigate to your data directory containing .raw/.mzML/.d.zip files
cd /path/to/your/data
# Ensure params.yml and dataset.csv exist
# Run the workflow
diann-snakemake --cores 8 -p all
# Or use snakemake directly
snakemake -s /path/to/Snakefile.DIANN3step.smk --cores 8 all
Docker Image
# Build the DIA-NN Docker image (default: 2.3.2)
docker build --platform linux/amd64 -f docker/Dockerfile.diann -t diann:2.3.2 .
# Build a specific version
docker build --platform linux/amd64 --build-arg DIANN_VERSION=2.3.1 -f docker/Dockerfile.diann -t diann:2.3.1 .
# Test the Docker wrapper
diann-docker --help
Architecture
Three-Stage DIA-NN Workflow
The core architecture is built around DIA-NN's three-stage processing:
Step A: Library Search (src/diann_runner/workflow.py)
- Input: FASTA database only (no raw files)
- Uses deep learning predictor to generate predicted spectral library
- Outputs:
{workunit_id}_predicted.specliband.config.json
Step B: Quantification with Refinement (src/diann_runner/workflow.py)
- Input: Predicted library + raw/mzML files (can be a subset)
- Generates refined empirical library from actual data
- Optionally generates quantification matrices (controlled by
quantifyparameter) - Uses
--reanalysefor match-between-runs (MBR) - Outputs:
{workunit_id}_refined.specliband.config.json
Step C: Final Quantification (src/diann_runner/workflow.py)
- Input: Refined library + raw/mzML files (can be different/larger set than Step B)
- Produces final quantification results
- Can reuse
.quantfiles from Step B with--use-quantflag - Outputs: Final TSV reports and matrices
Configuration State Management
Critical: The workflow uses .config.json files to ensure parameter consistency across stages.
- Each stage saves a
.config.jsonfile alongside its output (e.g.,predicted.speclib.config.json) - Steps B and C load the config from the previous step to ensure all parameters (var_mods, threads, qvalue, etc.) remain consistent
- This prevents common mistakes like changing modifications between stages
Implementation in src/diann_runner/workflow.py:
to_config_dict(): Serializes all workflow parameterssave_config(): Saves config JSON after each stagefrom_config_file(): Loads workflow from config
Module Structure
All source modules are located in src/diann_runner/:
src/diann_runner/workflow.py - Core workflow generation
DiannWorkflowclass: Manages all three stages with shared parameters_build_common_params(): Builds DIA-NN CLI arguments shared across stages_write_shell_script(): Generates executable bash scripts- Each
generate_step_*()method creates a bash script for that stage
src/diann_runner/diann_docker.py - Docker wrapper for DIA-NN
- Automatically detects Apple Silicon and uses
--platform linux/amd64 - Mounts current directory to
/workin container - Preserves UID/GID on Unix systems for correct file permissions
- Image and runtime come from
--image/--runtimearguments only; this wrapper reads no environment variables.
src/diann_runner/plotter.py - QC plotting utilities (diann-qc command)
src/diann_runner/qc_report.py - Markdown QC report generation (diann-qc-report command)
src/diann_runner/cleanup.py - Cleanup utilities (diann-cleanup command)
src/diann_runner/thermoraw_docker.py - Thermo .raw file conversion via Docker (thermoraw command)
src/diann_runner/snakemake_cli.py - Snakemake workflow runner (diann-snakemake command)
src/diann_runner/prolfquapp_docker.py - Prolfqua QC integration (prolfquapp-docker command)
src/diann_runner/prozor_diann.py - Protein inference CLI (prozor-diann command)
- Re-annotates DIA-NN report with protein IDs from FASTA using greedy parsimony
- Outputs
_prozor.parquetwith updated Protein.Ids and Protein.Group columns
src/diann_runner/prozor/ - Protein inference subpackage (Python port of R prozor)
ahocorasick.py: Backend abstraction for Aho-Corasick pattern matchingannotate.py: Peptide-to-protein annotation using multi-pattern matchingsparse_matrix.py: scipy.sparse matrix for peptide-protein relationshipsgreedy.py: Greedy parsimony algorithm for minimal protein set inference
contrib/oktoberfest/ - Optional Koina/Oktoberfest integration (see contrib/oktoberfest/README.md)
src/diann_runner/snakemake_helpers.py - Helper functions for Snakemake
detect_input_files(): Detects .d.zip, .raw, or .mzML files with priority logicparse_flat_params(): Transforms flat Bfabric executable keys to nested structureparse_var_mods_string(): Parses modification strings into tuplescreate_diann_workflow(): Factory function to initialize DiannWorkflow from parsed paramsget_final_quantification_outputs(): Returns output paths based on Step B vs Step C
Snakemake Workflow
The Snakefile.DIANN3step.smk orchestrates the complete pipeline:
- File conversion (
.raw→.mzMLor.d.zip→.d) - DIA-NN execution (generates and runs bash script)
- Prozor protein inference (re-annotates proteins from FASTA)
- QC report generation
- Results packaging and upload to bfabric
Key features:
- Reads configuration from
params.ymlin the working directory - Dynamically detects input file types (
.raw,.d.zip, or.mzML) viadetect_input_files() - Integrates with FGCZ infrastructure (bfabric, prolfqua)
- Uses Docker containers for msconvert and prolfqua
- Supports optional Step C (controlled by
enable_step_cparameter) - Runs prozor protein inference after DIA-NN quantification
Bfabric Parameter Flow Architecture
The workflow integrates with Bfabric LIMS, which requires a specific parameter transformation pipeline:
Parameter Flow:
Bfabric executable definition (executable_A386_DIANN_3.2.yaml)
→ GUI parameter selection
→ YAML with flat keys (params.yml)
→ Python nested structure
→ DiannWorkflow
→ DIA-NN CLI commands
Key Components:
-
Executable Definition (
bfabric_executable/executable_A386_DIANN_3.2.yaml)- Defines GUI parameters with flat keys like
06a_diann_mods_variable - Uses hierarchical numbering (06a, 06b, 06c) for logical grouping
- Parameter order affects GUI layout in Bfabric
- Upload with the Makefile beside it (
make validate,make upload ENV=TEST). Note thatuploadonly ever CREATES a new executable — seedocs/BFABRIC_DEPLOY.mdfor updating one in place.
- Defines GUI parameters with flat keys like
-
YAML Output (
params.yml)- Generated by Bfabric with flat keys matching the executable definition
- Example:
06a_diann_mods_variable: '--var-mods 1 --var-mod UniMod:35,15.994915,M'
-
Parsing Layer (
snakemake_helpers.py)parse_flat_params(): Transforms flat Bfabric keys to nested Python structureparse_var_mods_string(): Parses modification strings into tuples- Maps Bfabric keys to workflow parameters:
06a_diann_mods_variable→diann['var_mods']11b_diann_protein_relaxed_prot_inf→diann['relaxed_prot_inf']12a_diann_quantification_reanalyse→diann['reanalyse']
-
Snakefile Integration (
Snakefile.DIANN3step)- Calls
parse_flat_params(config_dict["params"])to transform parameters - Passes nested structure to
DiannWorkflowconstructor
- Calls
-
Workflow Generation (
src/diann_runner/workflow.py)DiannWorkflowclass uses nested parameters_build_common_params(): Converts to DIA-NN CLI flags- Conditionally adds flags based on boolean parameters:
--relaxed-prot-inf(ifrelaxed_prot_inf=True)--reanalyse(ifreanalyse=True)--no-norm(ifno_norm=True)
Important Rule: Complex Python code must ALWAYS go in snakemake_helpers.py, never directly in the Snakefile. The Snakefile should only orchestrate rules and call helper functions.
Important Patterns
No Print Statements for Logging
NEVER use print() for debugging or informational output. If logging is truly needed, use loguru. Most CLI tools should run silently - let the underlying tools (Snakemake, DIA-NN, etc.) produce their own output.
File Management Policy
NEVER use symlinks. Always use direct file references. This project policy prohibits symlinks in all scenarios - they add unnecessary complexity and can cause issues with some tools.
Fail Fast on Bad Config
NEVER use .get() with default values for required config parameters. If the YAML/config is malformed or missing a required key, the app should fail immediately with a clear KeyError. Silent defaults hide configuration errors and make debugging harder.
# WRONG - hides missing config
converter = WORKFLOW_PARAMS.get("raw_converter", "thermoraw")
# CORRECT - fails fast if key missing
converter = WORKFLOW_PARAMS["raw_converter"]
The Executable YAML Is the Source of Truth
bfabric_executable/executable_A386_DIANN_3.2.yaml defines the UI and is the single source of truth for parameter values, enumerations, and sentinel strings. Changes always flow in this direction:
executable YAML (source of truth)
→ Bfabric GUI → params.yml
→ parse_flat_params() in snakemake_helpers.py
→ DiannWorkflow
→ tests
When adding or changing a parameter:
- Define it in the executable YAML first (enumerations, default value, type)
- Add parsing in
parse_flat_params() - Wire it through
create_diann_workflow()if needed - Update tests to match the YAML values exactly (e.g.,
AUTOnotauto)
Sentinel values must be consistent: use AUTO (uppercase) for "auto-determine" parameters, matching the YAML. Tests must use the same strings the YAML defines — tests reflect the UI, not the other way around.
A pipeline_diann_version enumeration entry also needs a matching image key in src/diann_runner/config/defaults_server.yml under both images.docker and images.apptainer, or the run fails with a KeyError.
The same file is duplicated as slurmworker/config/A386_DIANN_23/executable_A386_DIANN23plus.yaml with no sync mechanism — diff and copy across after editing either.
Note: a definition pulled out of the Bfabric web GUI as an "XML Export" is not uploadable, and bfabric-cli executable upload only ever creates a new executable. See docs/BFABRIC_DEPLOY.md.
Flexible File Lists Between Stages
A key design feature is that Step B and Step C can use different file lists:
# Fast library building: use subset in B, all files in C
workflow.generate_all_scripts(
fasta_path='/path/to/db.fasta',
raw_files_step_b=['pilot1.mzML', 'pilot2.mzML'], # 2 files for fast library
raw_files_step_c=['s1.mzML', ..., 's50.mzML'], # All 50 files for quantification
quantify_step_b=False # Skip quantification in B, only build library
)
This pattern is used when:
- You have many files (50+) and want fast library building
- Building library from representative samples, then quantifying everything
- Running pilot → production workflows
Variable Modifications Format
Variable modifications use tuples of (unimod_id, mass_delta, residues):
var_mods = [
('35', '15.994915', 'M'), # Oxidation (Met)
('4', '57.021464', 'C'), # Carbamidomethyl (Cys)
('21', '79.966331', 'STY'), # Phospho (Ser/Thr/Tyr)
]
Default Binary: Docker Wrapper
By default, all workflow scripts use diann-docker as the DIA-NN binary. Override with:
--diann-binCLI flag"diann_bin"in config JSONdiann_binparameter in DiannWorkflow constructor
Key Output Files
DIA-NN 2.3+ outputs .parquet files by default. Downstream tools read the parquet directly.
out-DIANN_libA/
├── WU{id}_report-lib.predicted.speclib # Step A: Predicted library
└── WU{id}_libA.config.json
out-DIANN_quantB/
├── WU{id}_report-lib.parquet # Step B: Refined library (parquet format)
├── WU{id}_quantB.config.json
├── WU{id}_report.parquet # Main report (parquet) — native Run column
├── WU{id}_report.pg_matrix.tsv # Protein group matrix
├── WU{id}_report.stats.tsv # Statistics
└── diann_quantB.log.txt # DIA-NN run log
out-DIANN_quantC/ # Only if enable_step_c=True
├── WU{id}_report-lib.parquet # ★ Final library
├── WU{id}_report.parquet # ★ Main results (parquet) — native Run column
├── WU{id}_report.pg_matrix.tsv # ★ Protein matrix
├── WU{id}_report.stats.tsv # Statistics
└── diann_quantC.log.txt # DIA-NN run log
All downstream QC tools (diann-qc, prolfqua QC via prolfquapp ≥2.2.6, pmultiqc)
read the native WU{id}_report.parquet directly. The runner no longer emits a
Run→File.Name-renamed WU{id}_report.tsv (the old DIA-NN 1.x prolfqua shim).
Note on FASTA files:
- Step A requires FASTA input via
--fastafor library generation - Steps B and C can optionally use FASTA (via
fasta_fileparameter or DiannWorkflow constructor) for protein inference and annotation - FASTA files are NOT copied to output directories - only the original path is referenced
- The FASTA path can be stored in the
.config.jsonfiles for consistency across stages
Common Workflows
1. Create a Configuration File (Recommended First Step)
Before running workflows, create a reusable configuration file with your default parameters:
# Create config with your standard settings
diann-workflow create-config \
--output my_defaults.json \
--workunit-id WU123 \
--var-mods "35,15.994915,M" \
--threads 32 \
--qvalue 0.01
# View the generated config
cat my_defaults.json
This config can then be used with any workflow command via --config-defaults, with CLI arguments overriding config values as needed.
2. Standard Workflow: Same Files for B and C
Without config file:
diann-workflow all-stages \
--fasta /path/to/db.fasta \
--raw-files sample*.mzML \
--workunit-id WU123 \
--var-mods "35,15.994915,M" \
--threads 32
With config file (recommended):
# Using config defaults, only specify file-specific args
diann-workflow all-stages \
--config-defaults my_defaults.json \
--fasta /path/to/db.fasta \
--raw-files sample*.mzML
3. Fast Library Building: Subset for B, All Files for C
When you have many files (50+), use a subset for fast library building in Step B, then quantify all files in Step C:
# Step A: Library search (FASTA only, no raw files)
diann-workflow library-search \
--config-defaults my_defaults.json \
--fasta db.fasta
# Step B: Fast library refinement using subset (no quantification)
diann-workflow quantification-refinement \
--config out-DIANN_libA/WU123_predicted.speclib.config.json \
--predicted-lib out-DIANN_libA/WU123_predicted.speclib \
--raw-files pilot1.mzML pilot2.mzML \
--no-quantify # Skip quantification, only build refined library
# Step C: Full quantification with all files
diann-workflow final-quantification \
--config out-DIANN_quantB/WU123_refined.speclib.config.json \
--refined-lib out-DIANN_quantB/WU123_refined.speclib \
--raw-files sample*.mzML
Note: The --config parameter in Steps B and C points to the .config.json file from the previous step, ensuring parameter consistency across stages.
Documentation
Additional documentation is available in the docs/ directory:
docs/USAGE_EXAMPLES.md- Usage guide with quick reference and detailed patternsdocs/DIANN_PARAMETERS.md- Comprehensive DIA-NN parameter reference (compiled from GitHub repo, issues, and discussions)README_DEPLOYMENT.md- Deployment guide for production serverscontrib/oktoberfest/docs/- Koina/Oktoberfest integration (optional)
When troubleshooting DIA-NN issues: Consult docs/DIANN_PARAMETERS.md for parameter explanations, common issues, and links to relevant GitHub discussions.
Testing Notes
- Tests are in
tests/directory test_workflow.pytests the DiannWorkflow class- Run tests before committing changes to workflow generation logic
