Imported from pietrek928/nest (
AGENTS.md). Install upstream withnpx skills add pietrek928/nest. Copyright stays with the author.
Agent instructions (nest)
Standing rules for coding agents. Prefer this file over chat memory for repo conventions. Human docs: README.md. Deeper domain archive (not always-on): docs/agent-domain-notes.md.
Commands
# Python deps / rebuild native after C++ edits
uv pip install -e .
# Or targeted native rebuild
cmake --build build --target geometry graph
# Python tests
uv run pytest tests/ -q
# Type check (needs: uv sync --extra test)
uv run mypy
# C++ tests (configure once with -DNEST_GRAPH_BUILD_TESTS=ON)
cmake -S . -B build -DNEST_GRAPH_BUILD_TESTS=ON
cmake --build build --target geometry_cpp_tests graph_cpp_tests
# C++ warnings: -Wall -Wextra -Wpedantic on first-party targets (default).
# Strict: cmake -S . -B build -DNEST_GRAPH_WERROR=ON
Native sources are listed in tool.uv.cache-keys.
Code style
Python
- Python 3.12+: never add
from __future__ import …(includingannotations). - Prefer
X | None; quote forward refs only when required. - Imports only at module top. No mid-function / mid-branch
import/from … import(including insiderun_build_graph). Exception: true circular-import break with a one-line comment naming the cycle — prefer restructuring over lazy import. - No nested
defunless there is a concrete reason (closure over locals that cannot be args, or a one-shot callback required by an API). Prefer module-level / file-scope helpers. Do not nest helpers insiderun_build_graph/ propose hot paths for “locality.”
C++
- No custom namespaces (no anonymous
namespace {…}, no helper namespaces). Helpers at file scope in.ccorinlinein headers. - Namespace aliases for third-party APIs are fine (
namespace nb = nanobind;). - Avoid
staticon functions unless needed for linkage; prefer sharedinlinehelpers ininternal/internal.h.
Git
- Do not commit unless the user asks.
- Isolate with stash, not checkout. To temporarily drop WIP for a baseline/ablation bench, use
git stash push -m '<label>' -- <paths>(thengit stash pop/apply), notgit checkout HEAD -- <paths>orgit checkout -- .. Checkout discards uncommitted work; stash keeps it recoverable. Never path-checkout dirty propose/pack/geometry files “just to compare.”
Planning
Do this before locking a plan and before each implementation stage. Not a finishing checklist.
- Dedup / one gate. Grep for the predicate you are about to add (lex hold, round-4 keys, prepend/union, zone skip, gravity vector, mix niches, rim restore). The plan must cite the existing function. Do not add a second helper, skip site, or restore path. Permission (
ZONE_PROPOSERS) stays distinct from staging (packed_n/ void /use_*flags). One truncation/cut per pipeline. - Validate vs code. Check comments that already claim the behavior against the actual condition. Check plan claims against live helpers (
transform_row_key,lex_count_area_better,part_extents,placement_obstacles, …). Prefer extending a named path over a parallel RCL/beam/hold. - Stage + bench. Split the plan into letters. After each: smallest relevant
pytest;uv run python scripts/benchmark_pipeline.py --tags <case> --seeds 0 --propose shipped --gate; if propose/mix/nest changed,NEST_BUILD_GRAPH_ITERS=2 uv run python -m nest_graph.build_graph. Snapshot a baseline before the first letter. - Every letter: bench → conclude → improve. After each letter’s implementation, run the dual gate before anything else. Read letter telem + area/parts/time and write one conclusion (what bottleneck moved / what did not). If degraded, indep fail, or letter telem absent: stay on the letter — research the named hot path → one unify patch → re-bench until quality ≥ snapshot (or the user redirects). If green but telem still shows the bottleneck this letter was meant to move: one evidence-driven improvement patch on the same helper before locking (prefer gain, not floor-only). Do not start the next letter while the current letter’s expected telem is absent or quality is below snapshot.
- Bench conclusions drive the patch. The dual’s letter telem (and any branch table) names the next hypothesis; implement that one; re-bench. Opportunistic extras from a green dual are OK only when they extend the same named helper and keep indep + ≥ snapshot — never a parallel mechanism “while we’re here.”
- Bench the letter's component. Gate tags/cases must actually run the helper you changed (letter telem present, e.g.
niche_pos/contact_grg_upserts/pin_added/dfs_passes/refine_ms/history_expand/cluster_copy/free_space_cloud). Do not declare pass from a tag that never hit that path. To isolate vs other stages, mute unrelated existinguse_*/enable_*flags —uv run python scripts/benchmark_pipeline.py --tags <case> --seeds 0 --propose shipped --cfg enable_lns_rebuild=false enable_cluster_repack=false enable_gravity_compaction=false— and compare muted vs shipped on the same fixture. Prefixselection.for SelectionConfig (e.g.--cfg selection.dfs_passes=1). Do not add a second pack/credit/pin just to turn something off. Muted runs are compare-only; letter pass is still shipped vs snapshot. - Miss → improvement loop (same letter). Hard stop only if
independent_ok=false. Miss if quality < 0.9× best-so-far this run (and not below 0.9× last shipped bench for that tag), or time >1.5× with no quality gain, or the letter’s expected telem is absent. On miss: loop — research → patch → re-gate — until the letter passes or the user redirects. Do not start the next letter or pile a new parallel mechanism while looping. - Degradation → research loop (same letter). Any drop vs the letter’s snapshot / best-so-far (area, parts, or indep) is a fail to improve, not a soft OK. Even if still ≥0.9× the floor, do not lock the letter or move on while degraded: research (telem + named hot path) → one unify patch → re-bench until quality is ≥ snapshot (or the user redirects). Reverting a harmful patch counts as a loop step; “within noise / gate still green” does not excuse stopping below baseline when the task is to improve.
- Historical Qs stay in docs/agent-domain-notes.md. Do not lock one-off Q-numbers in this file.
Improvement loop (research + unify)
During each miss cycle, degradation cycle, and while a letter is still open:
- Stuck → research, then unify. Do not stack another special case. Read telem + the named hot-path functions; look for a better existing lever or a hybrid of two winners (shared helper, lex pick, soft scale). Optimize / simplify the current path before inventing a third.
- Do not alter the problem. Never make a gate easier by changing the fixture: no larger/smaller board, no catalog/demand/mix/iters retune, no lowered floors. Restore any such edits. Work the algorithm on the original case.
- Simplify while looping. Complex, duplicated, or deeply nested logic makes the next optimization impossible. If the path is getting hard to follow, stop adding features and unify/flatten first (one gate, one helper, delete dead flags).
- Validate logic as you go. Periodically check that comments, conditions, and call sites still match (hold vs override, colonize vs pin, floor vs actual predicate). Catch contradictions early — do not wait for the letter to finish.
- Research for improvements. Read telem + named functions on the hot path; form one hypothesis; one patch; re-bench. Prefer levers already in-tree (flags, seeds, budgets, existing helpers) before inventing a third path.
- Hybrid unify. If two solutions each win on different axes (density vs speed, swap-on vs swap-off, cascade vs free emit, …), do not keep both forever and do not pick one blindly. Look for a combined / unified form that keeps each advantage to the extent possible (lex pick, shared helper with both predicates, soft scale instead of hard skip, …). Cite both winners in the patch rationale.
- Unsure what hurts → telemetry first. If the failure mode is opaque, add the smallest bench/telem that names the stage (void props/graph/nest/refine, cascade stop, pin add, rim drop, …), re-run, then patch from evidence — not from guess stacks.
- Unify as you iterate. Every loop is also a cleanup pass: fold duplicates into one gate, flatten nested branches, delete dead flags.
- Keep logic clean and consistent. Same predicate → same helper; same SoT → same call site family; comments must match code. Prefer one readable path over clever special cases.
- Improve is the bar. Letter success is indep OK and quality ≥ snapshot (prefer gain). Telem-only / structural ships that leave area below the letter baseline stay in the degradation loop.
- Bench conclusions drive the next patch (see Planning). While a letter is open, the dual’s letter telem names the hypothesis — not a guess stack or a parallel mechanism.
Cross-track synergy
When a plan spans propose, post-pack, compose, DecisionGraph, and MCTS, tracks should complement through existing one-gates — not parallel stores, beams, or credit ledgers.
- One SoT per concern. Motif cross-iter → MotifBase pairs (
pattern_archive,upsert_from_contacts); same-iter N-way →ClusterPattern+merge_cluster_patterns; compose locks →sequential_accept_motif_cohorts→_nest_with_locks; refine restore →apply_refine_with_restore; MCTS cache →cheap_pack_cache_key. Extend the named path; grep before adding a second. - Downstream consumes upstream. Example chain: stamp/repack accept → pair upsert → next-iter
motif_patterns_for_inject→cluster_copy/ repair patterns → cohort beam →bind_epochMotifJoin → refine fracture. A later letter should not invent storage the earlier letter was meant to feed. - Hybrid over either-or when predicates differ. Clearance-valid stamp poses may fail contact upsert; archive pair relatives from the accepted pattern and masked contact upsert. Leader-star MotifJoin (k−1 edges) over full clique when refine budget matters. Score-sum tie in restore over count-only when C++ lex already uses score.
- Wire prerequisites before dependents. Extend cache keys before enabling new AMAF dimensions (
rule_id,cohort_sig). Prove M2member_hits/ compose telem before N-way MCTS macros. Conditional letters stay gated on telem from the prior letter — do not ship both blindly. - Orthogonal layers stack; duplicate gates don't.
pose_kindspatial bias andrule_idpreset selection are separate — OK to combine with telem. Two merge helpers, two motif archives, or two refine restore predicates are not. - Cross-track regressions need cross-track telem. Handoff keys include
repack_motif_upserts,motif_clique_*,motif_compose_*,member_hits,refine_score_accept,mcts_rule_id. A miss in track B after track A shipped often means A's output never reached B's gate — research the named hot path, don't add track C.
Open Q-table for multi-track work: docs/agent-domain-notes.md (net-only Q&A).
Nesting invariants
- Output must be collision-free (independent set). Transient DFS overlaps OK;
refine_selection/finalize_selectionmust not return overlaps to Python. - Default pipeline:
compose_and_nest_selection→nest_by_scores→ 3a (block_replacelock-swap, mid-pack) →refine_selection(unlocked) →finalize_selection→ 3b (contact-CC re-emit) → stamp fallback → post-pack gravity. Attract is finalize/tie-break only. Cheap expand:local_swap=Falseunless large_void (then dual lex on/off). Heavy leaf always dual (Q105: dual = heavy OR large_void). Natives reusemake_polygon_graph/NestState.native_geoms. DFS locks stay unset (finalize re-inserts clear locks).nest_by_graphremains forscore_rules/ tests only. - Do not add
NEST_DFS_MIN_COLLISIONS_*env vars; caps live inrefine_dfs.cc. - Board membership = locked void solids (pad complement + exterior slabs + sheet holes) ∪ packed parts — not
fully_inside/footprint_inside(oracle/tests only). - One obstacle assembler:
placement_obstacles(voids, packed). - Clearance SoT (same obstacles, different margins):
- emit →
emit_packing_clear(Penetrating, margin 0) - selection/polish →
is_pose_clear(Scene margin) - guidance →
valid_at/PlacementScene.is_valid(min_dist+ε)
- emit →
- Packing collide =
Geometry.intersects/ Penetrating only. Do not soft-filter in Python. - Contact: use
distance() <= threshold, never.buffer(gap).intersects(). Cluster:≤ 2·gap. Board adj: standoff≤ min_dist + 2·gap.
Propose / post-pack
ProposeContextis propose-only (emit/rank). Do not widen it or import it frombuild_graph/graph. Post-pack edits useSelectionEditCtx.- Emit order is static (not a registry). Poles →
pocket_fit→cluster_copyprecede sweepers. Funnel keys:round(x,y,θ, 4)=propose/void_selection.transform_row_key. - Mid-pack rim:
board_edgereserve beforeside_packkey claims; late kiss inlocal_se2(cached exterior ring). enable_gravity_compactiongateslocal_se2floater pole SE(2) toward an explicit void pole (default on). Ban corner / min-x+y sheet gravity (compact_selectiondeleted). Distinct from proposeborder_focus.cluster_relocate= rigid island ΔT (keep).local_se2= per-part SE(2).
Module map (propose)
| Module | Owns |
|---|---|
propose/types.py |
ProposeContext, extras, make_propose_context |
propose/pipeline.py |
staged _collect_candidates, ranking handoff |
propose/transform_batch.py |
mix / window / angle project / graph-valid carry |
propose/heavy_polish.py |
polish budget + DFS dispatcher (apply_dfs_refinement) |
propose/telem.py |
void_leak assembly (assemble_void_leak) |
propose/selection_edit.py |
SelectionEditCtx for local_se2 / repack / relocate |
propose/post_pack.py |
repack → relocate → local_se2; prepare_post_pack / apply_post_pack_and_telem |
propose/block_replace.py |
3a cohort lock-swap; 3b hole re-nest (maybe_block_hole_renest) |
propose/placement_common.py |
placement_obstacles, is_pose_clear, independence helper |
pack/ctx.py |
PackIterCtx, RefinePackBox, stage result types |
pack/stages.py |
compose/refine, mid_pack, first_pass, post_pack, rim_before |
pack/credit.py |
run_void_leak_and_niche_credit, finalize_iter_mcts |
pack/cheap.py |
pack_execute_snapshot, cheap cache key + compose/refine adapters |
pack/execute.py |
record_outer_iter_expand, execute_pack, MCTS multi-sim |
rules/evolve.py |
improve_rules, dedupe_rule_sets, demo/seed rule factories |
pack/ |
Macro-MCTS orchestration (cheap expand vs best-leaf polish) |
graph/pose/pose_graph.h |
Pose MIS (replaces ElemGraph) |
graph/decision_arena.h / motif_base.h / se2.h / contact_relation.h |
C++ arena + BoardSnapshot + MacroNicheArchive, motifs, SE2, ContactGRG+GCI |
graph/decision/mcts_agent.h |
MctsAgent, leaf_reward, path_reward_beats (UCB1/PW/AMAF) |
graph/decision/action_gen.h |
generate_macros, region_to_zone / zone_to_region |
geometry/common/nfp_lite.h |
Motif inject polish (nfp_lite_relative); not a MotifBase miner |
build_graph owns graph/selection + Macro-MCTS outer loop (cheap expand vs best-leaf polish). Do not re-export moved names from build_graph.
Hard bans
- No Python grids / Shapely fallbacks when C++ polish or Scene fails — fix the C++.
- Hot paths:
nest_graph.geometry.Geometry+NestState.native_geoms. Do not loopGeometry.from_shapely()in propose emit / tightness / coverage. - Prefer raw ring samples + batch
valid_at; do notbuffer(-min_dist)seed erosion. Ribbon annuli OK for frontier focus. - Do not add Clipper2, libnest2d, or C++ Voronoi.
- Do not OR guidance
valid_atwith packing independence. - Void-fill last resorts: no corner gravity; PSO only if
props_pole≈0after OOS-1+4. B2: no void-aware nest/refine MIS until B0/B1 exhausted; rim drop >2% = fail.
Glossary (confusion traps)
| Term | Meaning |
|---|---|
border_focus |
Propose rim/kiss zone — not post-refine gravity |
void_seek |
Large free void path; may override packed_near_border |
emit_packing_clear |
Emit packing SoT (Penetrating); allows rim Touch |
is_pose_clear |
Selection/polish Scene clearance |
valid_at |
Guidance clearance only |
board_edge reserve |
Claims rim keys before side_pack / explorers |
Verify before finishing
- Run the smallest relevant
pytestfor touched code. - After C++ edits: rebuild (
uv pip install -e .or cmake targets above), then run matching C++ or binding tests. - Confirm selection outputs stay pairwise independent when changing pack/polish paths.
- Staged impl follows Planning (gate after each letter; on miss run the improvement loop, including unify/telem before the next letter).
