Imported from yhyu13/AlphaGOZero-python-tensorflow (
new/AGENTS.md). Install upstream withnpx skills add yhyu13/AlphaGOZero-python-tensorflow --skill new. Copyright stays with the author.
AGENTS.md — new/ (PyTorch AGZ port)
Repo-specific notes for future Kilo sessions working on this folder.
Scope
This is a clean, modern, runnable PyTorch reimplementation of the AGZ pipeline. The legacy TF 1.x code in the parent folder is not part of this; do not import from it. The interface contract is "self-play → replay buffer → train → evaluate against best", mirroring the original but without its many quirks.
Stack
- Python 3.10+, PyTorch 2.x. No TF.
numpyfor Go state,sgffor any future SGF work.requirements.txtpins:torch>=2.0,numpy,sgf>=0.5. Add more deps with care.
Layout
config.py FLAGS + HPS (lazy: parses on first attribute access)
main.py CLI dispatch — one function per mode
game/go.py Position, rules, KO, captures, area scoring
game/features.py 17-plane encoder + 8 symmetries
model/network.py AGZNet — pre-activation ResNet + dual heads
model/resnet.py PreActResidualBlock
mcts/mcts.py PUCT MCTS, Dirichlet noise at root
training/
dataset.py ReplayBuffer (FIFO, fixed-size)
selfplay.py Self-play game generator
train.py Train step (policy + value loss + symmetry aug)
evaluate.py Head-to-head evaluator
tests/
test_smoke.py unittest suite (run with `python -m unittest new.tests.test_smoke`)
Conventions
- Config goes in
config.py.FLAGSandHPSare populated lazily via__getattr__so importing this module never consumes the test runner's argv. New hyperparameters go inFLAGS(argparse) or propagate viaHPS. - No global state.
NetworkandMCTSare pure functions of their inputs. Passdeviceexplicitly. - Symmetry augmentation lives in
game.features.apply_symmetry. Use_mirror_policyintraining/train.pyfor the policy target — keep both in sync. - Replay buffer samples carry
_winneras a side-channel attribute (set byattach_winner). The value target is+1ifsample.player == sample._winner, else-1. - Play index 0 is black.
STONE_BLACK = 1,STONE_WHITE = 2,STONE_EMPTY = 0. Pass is(-1, -1). - MCTS root priors: Dirichlet noise on during self-play, off during evaluation. Controlled by
dirichlet_alpha+dirichlet_epsilon— passing both 0 disables noise. - Board size is configurable.
model/network.pyacceptsboard_sizeand the heads usen * nflatten sizes. The CLI defaults to 19; the smoke tests use 9. - Import policy.
training/__init__.pyonly re-exportsReplayBufferandSample(fromdataset.py). Other training modules must be imported explicitly (from new.training.selfplay import ...) to avoid chain-loadingconfig.py.
Quirks
Position._remove_groupuses negative markers (-color) so_restore_groupcan find the cells during undo. After capture is committed, cells are normalized back toSTONE_EMPTY(0). Don't introduce a separate "removed" set — the marker is the cleanest signal.- The MCTS scratch copy (
scratch = position.copy()) is mutated during selection. This is the only placePosition.copy()is hot; it's a numpy copy. extract_featureswalksposition.historyand replays it on a freshPosition. This is O(moves) per call — fine for batch sizes of 64, expensive for batches > 512. If you hit a wall here, cache(features, position)in the replay buffer.apply_symmetryand_mirror_policymust agree on how sym=0..7 maps to {rotations × reflections}. Look at both before changing either.Position.score()returns the area-score winner regardless of terminal state.Position.winner()requiresis_terminal()to be True — usescore()for truncated games (e.g. when self-play hitsmax_moves).config.pyuses module-level__getattr__soFLAGSandHPSare only parsed on first access. If you ever add eager top-level code toconfig.pythat touchesFLAGS, you'll reintroduce thepython -m unittestbug.
Verifying changes
# Smoke tests (no GPU required).
python -m unittest new.tests.test_smoke
# Tiny end-to-end smoke through the CLI.
python new/main.py --mode=selfplay --n_games_per_iter=1 --n_simulations=5 --n_resid_blocks=1 --n_filters=16 --n_moves_per_game=10
python new/main.py --mode=train --n_train_steps=3 --batch_size=4 --n_resid_blocks=1 --n_filters=16
python new/main.py --mode=evaluate --n_eval_games=1 --n_simulations=5 --n_resid_blocks=1 --n_filters=16
Tests
new/tests/test_smoke.py covers: Go rules (capture, alternation, marker cleanup), feature extraction (shape, symmetry round-trip), network (forward shape, no NaN), MCTS (visit policy sums to 1), replay buffer (add / sample / eviction). Run them with python -m unittest new.tests.test_smoke.
Things future agents will get wrong
- Reading
STONE_BLACK = 1and assumingSTONE_BLACKis "true" /1in arithmetic. It's just an enum. - Trying to use the legacy
model/from the parent folder. Don't. This is its own clean implementation. - Trying to "improve"
apply_symmetryor_mirror_policywithout re-deriving both. They are coupled. - Adding TF 1.x code paths. This folder is PyTorch-only.
- Running the default config (
19residual blocks,800simulations) on a CPU. Scale down first. - Adding
from new.config import FLAGStotraining/__init__.py. That re-exports trigger chain-loading and breakpython -m unittest.
