Imported from zhubonan/airsspy (
AGENTS.md). Install upstream withnpx skills add zhubonan/airsspy. Copyright stays with the author.
AGENTS.md
This file provides guidance to coding agents when working with code in this repository.
Venv for development is at .venv, activate with source .venv/bin/activate before any command. Always use uv pip install instead of pip install for installing packages.
Worktree virtualenv bootstrap
Worktree checkouts may start with an empty or partial .venv. When that
happens, mirror the dependency versions from the main repository instead of
letting uv resolve a fresh environment. This is especially important for torch:
the worktree must use the same torch build as the main repo.
Recommended workflow:
# From the worktree checkout
uv pip freeze --python /home/bonan/appdir/airsspy/.venv/bin/python > /tmp/airsspy-main-freeze.txt
sed 's#^-e file:///home/bonan/appdir/airsspy$#-e file://'"$(pwd)"'#' \
/tmp/airsspy-main-freeze.txt > /tmp/airsspy-worktree-freeze.txt
uv venv --python /home/bonan/appdir/airsspy/.venv/bin/python --clear .venv
uv pip install --python .venv/bin/python --no-deps -r /tmp/airsspy-worktree-freeze.txt
If running inside a sandboxed worktree, run the final uv pip install outside
the sandbox when necessary so uv can use its normal cache. Verify torch before
running ML-related tests:
/home/bonan/appdir/airsspy/.venv/bin/python -c "import torch; print(torch.__version__)"
.venv/bin/python -c "import torch; print(torch.__version__)"
Project Overview
airsspy is a Python library providing an ASE-based interface for Ab initio Random Structure Searching (AIRSS). It wraps the external buildcell executable (not bundled) and lets users construct search seeds programmatically, generate random structures, parse .res output files, run distributed searches via jobflow, and analyse results. Licensed under GPLv2.
Build, Install, and Test Commands
# Install in editable mode with dev dependencies (includes test tools)
uv pip install -e ".[dev]"
# Install with docs dependencies
uv pip install -e ".[docs]"
# Run the full test suite
pytest tests/ -q
# Run a single test file
pytest tests/test_seed.py
# Run a specific test
pytest tests/test_seed.py::test_some_function
# Lint and format (ruff, configured for line-length=88, target py39)
ruff check src/
ruff format src/
# Type check
mypy src/ --ignore-missing-imports
# Run all pre-commit hooks
pre-commit run --all-files
Architecture
Core modules (src/airsspy/)
-
seed.py—SeedAtoms(extendsase.Atoms) holds aBuildcellParaminstance (gentags) for cell-level buildcell parameters and per-atomSeedAtomTagentries stored in theatom_gentagsarray. Tag descriptors (BoolTag,GenericTag,RangeTag,NestedRangeTag) map Python attribute assignments to buildcell keyword syntax.SeedAtoms.write_seed()serializes to a.cellfile thatbuildcellreads. -
build.py—Buildcellwraps the externalbuildcellbinary viasubprocess.Popen.Buildcell.generate()pipes a seed's.cellcontent tobuildcelland parses the output back into an ASEAtomsobject. -
restools.py—RESFileclass and helper functions for parsing CASTEP.resfiles (TITL/CELL/SFAC blocks), extracting energies/volumes/spacegroups, and converting to ASE or pymatgen structures. Supportsfrom_packed()for concatenated RES files,from_file(),from_string(), and round-trip viato_res_lines(). -
common.py— DefinesBuildcellErrorandRelaxErrorexceptions. -
log.py— Shared CLI logging setup. -
utils.py— General utilities (string parsing, file pattern matching, k-point calculation). -
search.py— Reusable local-search helpers: formula sampling/injection for buildcell seeds, charge-neutral formula filtering, target-volume handling, and RSS post-relax pruning based on energy-per-atom and fingerprints.
Code-specific tools (src/airsspy/)
-
casteptools.py— CASTEP output parsing and file management:parse_dot_castep()(geometry convergence),parse_param(),RASH_prepare_seed(),extract_REM(),extract_result(),write_converge(),push_cell(),castep_finish_ok(),gulp_relax_finish_ok(). Exception classes:CastepRunError,CastepSkip,CastepManualTimedout. -
abacustools.py— ABACUS utilities: parses ABACUS running logs and STRU files, converts CASTEP.cellcontent to ABACUSSTRU, detects output logs, extracts REM metadata, and composes AIRSS-compatible result documents /.resoutput. -
scftools.py—SCFInfoclass for parsing SCF convergence data from.castepfiles. UsesScfDatanamedtuples per geometry step. Providesget_summary()statistics andplot_scf()/plot_conv()matplotlib plotting methods. -
gulptools.py— GULP utilities:geom_opt_progress()parses geometry optimisation,check_gulp()monitors for divergence/overflow,guarded_gulp()runs GULP with automatic termination. -
fullrelax.py—FullRelaxclass implementing self-consistent CASTEP relaxation with restart capability, alternating cell constraints, and state persistence. Includesparse_geom_text_output()for.geomfiles andgeom_to_cell()for converting last configuration to cell blocks. -
tools/modcell.py—replace_block()andmodify_cell()for programmatically modifying CASTEP cell files.
Scheduler (src/airsspy/)
scheduler.py— Unified scheduler interface:Schedulerbase class,Slurm,SGE, andDummyimplementations.Scheduler.get_scheduler()factory auto-detects the current environment.
Analysis sub-package (src/airsspy/analysis/)
-
collect.py— DataFrame collection utilities:collect_res_in_df(),combine_res_cryan()(cryan CLI wrapper),read_ca()(ca command parser),read_stream(),get_minsep_range(),get_entry()(ComputedEntry factory),export_dataframe_as_res(),get_pressure_gpa(). -
hull.py— Phase diagram plotting:PlotlyPDPlotterextends pymatgen'sPDPlotterwith Plotly backend for ternary/binary convex hulls. -
query.py— Bridge from jobflow store to analysis:collect_results_df()collects results from aSearchStoreinto an analysis-ready DataFrame.
Jobflow integration (src/airsspy/jf/)
-
documents.py— Pydantic output models:AirssJobDoc(one per job, contains N results) andAirssResultDoc(per-structure data).RelaxOutcomeenum for relaxation status. Both makers produce the sameAirssJobDoctype. -
runners.py— Pure computation classes:AirssCastepRelaxRunner,AirssCastepSinglePointRunner,AirssGulpRelaxRunner,AirssPp3RelaxRunner,AirssAbacusRelaxRunner,AirssAbacusSinglePointRunner,AirssScriptRelaxRunner,run_buildcell(),clean_files(), andcompose_task_doc(). Usable standalone or within Makers. -
ml_runners.py— Optional machine-learning potential runners.AirssMlRelaxRunnerandAirssMlSinglePointRunnersupport ASE-compatible calculators and optionaltorch_simbatch backends;compose_ml_task_doc()writes.resoutput from extxyz results with energy/forces/stress metadata. -
jobs.py— Jobflow Makers:AirssSearchMaker(build+relax N structures per job),AirssRelaxMaker(relax N provided structures),AirssValidateMaker. All produceAirssJobDocoutput. -
store.py—SearchStorequery layer wrapping maggmaMongoStore. Methods:retrieve_project(),retrieve_project_df(),list_projects(),list_seeds(),show_struct_counts(),throughput_summary().
CLI (src/airsspy/cli/)
Entry point: ap command (registered in pyproject.toml).
main.py— Top-levelairssClick group registered as theapentry point. Global options include--db-host,--db-port,--db-name,-v/--verbose, and-q/--quiet.cmd_deploy.py—deploy searchanddeploy relaxcommands (both with--dryrun).cmd_db.py—db list-projects,db list-seeds,db summary,db throughput,db retrieve-project.cmd_check.py—check airss,check scheduler,check database.cmd_tools.py—tools modcell.cmd_run.py— Local non-jobflow AIRSS runner, similar in role toairss.pl. Providesrun search,run relax, andrun sp; search supports CASTEP/GULP/PP3/ABACUS, while relax supports CASTEP/GULP/PP3/ABACUS/ML and single-point supports CASTEP/ABACUS/ML. Also supports build-only mode, packing results, MPI launcher options, walltime-buffer stopping, formula sampling (--formula,--elements,--target-volume, oxidation states), and RSS pruning (--prune*options).cmd_rank.py—rankcommand for ranking structures by enthalpy per formula unit. Reads.res(packed) or extxyz from stdin/file args. Core options:-t/--top,-de/--delta-e,-f/--formula(exact formula, comma-separated elements likeSi,O, or glob likeSi*),-nr/--absolute,-l/--long-labels,-s/--summary,-u/--unite(fingerprint merging),--input-format,--fingerprint-cutoff(default 10.0 Å),--unite-no-zweight(Z-weighting is on by default). Pipeline controls:--filter-name(glob on label),-p/--pressure(external pressure in GPa, applies PV correction),--unite-ethresh(energy threshold in eV/atom for pre-merge filtering, default 0.1),--unite-output(directory for merged groups),--unite-output-format(resorextxyz).cmd_pack.py—packandunpackcommands for text-based file concatenation and splitting.packconcatenates files (glob args or--from-dir);unpacksplits a packed file into a directory. RES unpack splits on TITL→END blocks; extxyz unpack splits per structure.--formatflag controls output format (resorextxyz).cmd_convert.py—convertcommand for RES ↔ extxyz conversion. Bulk:ap convert packed.res output.xyzorap convert input.xyz output_dir/. Single:ap convert -l Si-002 packed.res Si-002.res.
Ranking module (src/airsspy/ranking.py)
Standalone ranking logic with no pandas dependency. Key components:
StructureRecorddataclass — lightweight record with TITL metadata + species counts. Properties:reduced_formula(Hill system with cryan ordering: C first, H second if C present, O always last, rest by atomic number),n_formula_units(GCD of species counts),enthalpy_per_fu,volume_per_fu. Private_merged_peersfield tracks structures merged into this record byeliminate_similar()(used for--unite-output)._parse_res_fast()— parses only TITL + species counts, avoids pymatgen overhead.rank_structures()— groups by formula, computes relative energies, filters by delta_e.summary_structures()— most stable per composition (cryan-sequivalent).apply_external_pressure(records, pressure_gpa)— adds PV correction to enthalpy (H = E + PV), enabling ranking at non-zero external pressure.filter_by_name(records, pattern)— glob-based filter on structure labels (TITL label field).filter_by_formula(records, formula)— formula filter supporting three modes: exact reduced formula match, comma-separated element set (e.g.Si,Omatches any composition containing only Si and O), or glob pattern on reduced formula string.prefilter_records(records, ethresh)— removes structures whose energy per atom is more thanethresheV/atom above the best in their composition group. Applied before merge to avoid expensive fingerprint computation on clearly unstable structures._compute_distance_fingerprint()— computes sorted distance fingerprint using pymatgen'sget_all_neighbors(cutoff)(all periodic images within cutoff). Returnsnp.ndarray | None. Optional Z-weighting scales each distance asd * zmax^2 / (Z_i * Z_j)to distinguish atom-type pairs. Z-weighting is now enabled by default.eliminate_similar()— fingerprint-based merging (cryan-uequivalent). Accepts acutoffparameter (default 4.0 A) controlling the neighbour search radius. Uses numpy-vectorised fingerprint comparison. Records which structures were merged via_merged_peerson surviving records.format_header()/format_rank_line()— cryan-compatible tabular output.
Ranking pipeline order
The ap rank command processes structures through a fixed pipeline. The order matters because each stage reduces the work for subsequent stages:
- Read — parse input (packed
.resor extxyz) - Apply pressure (
-p) — add PV correction if external pressure specified - Filter —
--filter-name(glob on label) and-f/--formula(composition filter) - Pre-rank (
--unite-ethresh) —prefilter_records()removes high-energy outliers per composition - Merge (
-u) —eliminate_similar()merges fingerprint-duplicate structures - Post-rank (
-de,-t) —rank_structures()computes relative energies and applies final delta_e/top-N filtering - Display — tabular output, summary, or file export
Conversion module (src/airsspy/convert.py)
Lossless RES ↔ extxyz conversion with force support:
res_to_extxyz()— packed .res → single extxyz. Forces from atom columns 8-10 →SinglePointCalculator.extxyz_to_res()— extxyz → unpacked individual .res files (one per structure, named<label>.res).extract_structure()— extract single structure by label from packed .res or extxyz. Fast TITL-only scan first, then loads only the match. Output format from extension._atoms_to_res_lines()— builds RES lines from ASE Atoms, appends force columnsfx fy fzif present._parse_res_forces()— reads force columns 8-10 from .res atom lines.
Force column placement in .res: Symbol index x y z occ [spin] [fx fy fz]. Safe because cryan reads only columns 1-5, cabal reads up to column 7 (spin).
Search helpers (src/airsspy/search.py)
Local AIRSS search support outside jobflow:
- Formula sampling —
FormulaSamplingOptions,FormulaSamplingContext,build_formula_sampling_context(), andmake_seed_text_transform()inject sampled#FORMULA/#VARVOLdirectives into seed text. Formula pools can come from explicit formulas or enumerated element coefficient combinations and can be filtered by seed#NATOM/#NFORMconstraints and oxidation-state charge neutrality. - RSS pruning —
RssPruneOptions,RssCandidate,candidate_from_res(),pool_statistics(),should_flush_prune_pool(),select_pruned_candidates(), andprune_relaxed_pool()retain low-energy, fingerprint-unique candidates from relaxed pools before final output.
ABACUS and ML support
- ABACUS integration lives in both
abacustools.pyandjf/runners.py. ABACUS runners expect an ABACUS input file suffix fromcmd_deploy.SUFFIX_MAP, generate/consumeSTRUand ABACUS output directories, then compose AIRSS-style.resfiles. - ML support is optional under the
[ml]extra (torch-sim). The default ML path uses dynamic ASE calculator specs such asmodule.path:ClassName@model; whentorch_simis installed,ml_runners.pyhas batch relax/static helpers for model specs likemace:medium.
Reference implementations
The AIRSS reference Fortran code lives at ~/appdir/airss-git/:
internal/cryan/src/cryan.f90— Structure ranking/analysis tool. Key routines:read_res()(line 293),rank()(line 1667),compositions()(line 852),summary(),eliminate()(line 2200). Reads atom lines assymbol, index, x, y, zonly (ignores columns 6+). Formula ordering useselements_alphawith O→huge, C→0, H→0.1 if C present.internal/cabal/src/cabal.f90— Structure manipulation tool.read_res()(line 557) andwrite_res()(line 664) show full RES read/write. Atom format:Symbol index x y z occ [spin], with write formata4,i4,3f17.13,f4.1[,f7.2]. TITL format:label P V H spin spin_abs [dos] nat (symm) n - copies.
Key design patterns
SeedAtomsextendsase.Atoms; itsbuildproperty returns theBuildcellParamobject for setting cell-level tags likevarvol,symmops,numat. Per-atom tags (e.g.,posamp,tagname) are accessed viaseed.atom_tags[i].- The tag descriptor system uses Python descriptors (
__set__/__get__) that validate and serialize values into the buildcell cell-file format. Buildcellrequires the AIRSSbuildcellexecutable on$PATH. Tests that callbuildcellwill fail if it is not installed.- Multi-structure jobs: Both
AirssSearchMakerandAirssRelaxMakerproduce a singleAirssJobDoccontaining NAirssResultDocentries, reducing JobStore document count for high-throughput searches. - Runners are jobflow-free:
AirssCastepRelaxRunnerandrun_buildcell()can be used standalone without jobflow orchestration. - Local CLI search uses the shared runner layer rather than jobflow. It writes
.cell, relaxation outputs, and.resfiles in the selected work directory and honors astopfile plus scheduler walltime checks.
Dependency tiers
# Core (always installed)
ase>=3.17, castepinput>=0.1, click>=8.0, jobflow>=0.1, maggma>=0.50, numpy>=1.20,
pandas>=1.3.0, plotly, pymatgen>=2022.0.0, spglib>=1.16, tabulate, tqdm
# Optional
[dev] pytest, pytest-cov, ruff, mypy, pre-commit, twine
[ml] torch-sim
[docs] sphinx, pydata-sphinx-theme, myst-nb, sphinx-autodoc2, sphinx-design, etc.
Python >=3.9.
Testing
Tests live in tests/ with conftest.py providing fixtures (al_atoms, tmpfile). pytest is configured in pyproject.toml with --strict-markers --strict-config; the e2e marker is reserved for tests that run real CASTEP/ABACUS executables. Tests that depend on external programs such as buildcell, CASTEP, ABACUS, GULP, PP3, or optional ML backends should be skipped/guarded when the executable or package is unavailable. The suite covers core seed/build/RES utilities, CASTEP/GULP/ABACUS tools, ranking/conversion, search helpers, scheduler, jobflow documents/jobs/runners/store, ML runner helpers, and CLI commands (using click.testing.CliRunner).
Mandatory final review subagent
After implementing changes and before final response:
- Spawn a read-only reviewer subagent.
- Ask it to review the current branch diff against the base branch.
- The reviewer must focus on:
- correctness
- regressions
- missing tests
- race conditions
- security issues
- unintended unrelated edits
- Wait for the reviewer result.
- Fix any high-confidence issue.
- Summarize the review result in the final answer.
