Imported from zty1017/autosolver (
AGENTS.md). Install upstream withnpx skills add zty1017/autosolver. Copyright stays with the author.
AGENTS.md — AutoSolver AI 行为入口
本文件是所有 AI Agent 在本仓库工作的最高优先级项目指令。它只承载行为约束和读文档顺序;详细设计以 docs/ 下的新文档为准。
1. Source of Truth
当前事实源只包括:
AGENTS.mddocs/README.mddocs/next-agent-handoff.mddocs/current-state.mddocs/worldview.mddocs/problem-and-scoring.mddocs/target-architecture.mddocs/agent-governance.mddocs/evaluation-and-audit.mddocs/evolution-design.mddocs/distillation-and-submission.mddocs/implementation-roadmap.mddocs/acceptance-gates.mddocs/glossary.md
如果上述某个文件尚未存在,Agent 应先承认缺失,并在用户允许的范围内补齐文档,而不是从旧文档推断当前事实。
2. Legacy Archive Rule
docs/archive/** 是历史资料,不是当前事实源。
- 默认禁止主动读取
docs/archive/**来指导当前实现。 - 只有当用户明确要求“历史追溯 / 对比旧设计 / 查旧资料”时,才可读取归档内容。
- 读取归档内容后,必须明确标注其为历史材料,不得把其中描述当作当前设计或已验证事实。
3. Project Mission
AutoSolver 的目标排序:
- 产出尽可能强的 final Solver。
- 建立可信 AutoSolver 系统和证据链。
- 形成可读的项目报告与叙事。
质量优先,但禁止低信息量试错。剩余时间和 Online Submission 次数有限,每次运行、distill、提交都必须服务于明确假设、质量提升或不确定性降低。
4. Core Worldview
项目采用分层认知:
Online Evaluation
-> Scoring Model
-> Evolution Fitness / Island Objective
-> Solver Evolution
-> Distilled Solver
- Online Evaluation 是外部黑盒检验,但反馈稀缺。
- Scoring Model 已通过多轮 Online Submission feedback 校验;至少在已观测反馈和 public anchor
large_seed301上未观测到线上/线下偏差。 - Scoring Model 可作为当前 evolution 的 ground truth proxy,但不能假定完全覆盖 Hidden Evaluation Regime。
- Evolution Fitness / Island Objective 是驱动本地进化的选择压力,不等同于 Scoring Model。
- Online feedback 不应直接塞入 Mutator prompt;它应先用于 calibration、Reference Portfolio 和 Evolution Fitness 维护。
5. Target Architecture
主线是“从自由探索到结构化收敛”。
探索期:
- 以完整
solve()自由进化为主。 - 目标是让 Agent 自主发现策略族、失败模式和有效结构。
- MAP-Elites、Island Objective、Reflection、Evolution Insight、Adversarial Instance diagnostic 都服务于发现和解释。
收敛期:
- 从有效 Individual 中提炼
scaffold + strategy modules。 - 主进化转向模块化策略,完整
solve()自由进化只保留为低比例探索通道。 - Distilled Solver 可以是单策略,也可以是内部 Portfolio / router,但对外仍是一个 Solver。
6. Status Labels
所有机制、事实、设计必须使用以下状态之一:
已验证:有代码、数据或 Online Evaluation feedback 支撑。当前实现:代码里存在,但设计正确性和效果仍需审计。目标架构:希望达到的最终设计。实验机制:可探索,但不能作为默认主路径。已废弃:明确不再采用,只能作为历史背景。未知待验证:不能作为决策依据,必须先验证。
禁止把 目标架构、实验机制 或归档旧文档描述成 已验证。
7. Hard Constraints
- Agent 系统运行环境:
uv管理的 Python 3.12+。 - Solver 线上环境:Python <= 3.6,纯标准库,10 秒 / case。
- 本地 sandbox 执行 Solver 时应使用 Python 3.6;不得用 Python 3.12+ 结果替代线上兼容性判断。
- Solver 接口固定:
def solve(input_text: str) -> list。 - 输出必须是 Assignment 列表:
[task_id_list_str, [courier_id, ...]]。 - Courier 全局唯一;Task 不得重复覆盖;允许拒单。
.env等敏感文件不得写入文档或日志。- Online Submission 只能由人类手动执行;Agent 不得假装已提交。
8. AI Behavior Rules
- 每次只向用户抛出一个需要确认的问题。
- 不要启动无边界、无预算、无 post-run audit 的长 evolution;只有
acceptance-gates.md中对应 gate 已通过后,才允许受控扩大 run。 - 新的受控 evolution / structured search 前必须运行
python scripts/check_pre_run_gate.py;生成或更新 Gate 8 submission package 后必须运行python scripts/check_submission_package.py <package_dir>;run 结束后必须用 Candidate Audit、scripts/check_gate7_run.py或scripts/check_post_run_artifact.py中适用的 gate 固化结果,并回写 Source of Truth。 - 不要信任历史 run、summary、report、distilled solver,除非经过 candidate-level audit。
- 当前代码实现整体低置信;每个模块只能按 gate、审计报告和最新 run artifacts 逐项提升可信度,不得因为局部短跑通过就默认整个 evolution loop 正确。
- 不要因为某个局部指标变好就建议提交;必须比较 Reference Portfolio 和 full/profile matrix。
- 不要把人工强 Solver 描述为 AutoSolver 的主要成果;它们只能作为 Reference、baseline、gate 和 calibration 素材。
- 不要把外部 AI 编码工具的人工判断描述为 AutoSolver 的 Distill。Distill 是 AutoSolver Distiller 的系统行为,必须由流程调用和 report 证明。
- 不要把外部 AI 编码工具手工总结的策略描述成 AutoSolver 的
strategy module成果;structured strategy search 必须由项目脚本生成候选,并留下 strategy manifest、candidate audit 和 Reference comparison。 - 人工 insight、线上反馈、外部论文或附件 teacher 只能作为 hypothesis / operator spec / proxy diagnostic 输入;不得直接注入最终 solver、population parent、Reference Portfolio 或 Distill 结果。若要采用这些 insight,必须转成项目内可搜索的 strategy field 或 operator,由 AutoSolver 脚本生成候选并通过 gate。
- 用户已允许调用外部 LLM API 进行真实场景测试和自主推进;每次真实 run 仍必须固定预算、run_id、停止条件,并在结束后执行 Candidate Audit、Gate checker 和文档回写。
- 遇到旧实现与新文档冲突时,以新文档为准,并记录需要修复的实现差距。
- 如果短 Evolution run 已暴露 invalid mutation、parse fallback 或明显 full/profile 退化,下一步必须先修 operator / validator / accept gate,不得用更大 run 掩盖问题。
- Evolution operator attempt 允许有受控失败;目标不是零失败,而是高吞吐、可计量、可熔断、不可污染种群。
- hard failure 包括 parse/static/sandbox invalid、fallback-only output、exact duplicate;这些不得进入 Archive、parent selection、Reflection success、Insight success 或 Distill 候选。
- Reference-gate rejected 不是自动丢弃:如果它是 valid、非重复、具有明确 novelty / strategy value,可以进入隔离的
Exploration Pool/Novelty Archive,但必须标注为 novelty,不得冒充 quality improvement、submission candidate 或 Distill success。 - novelty Individual 可低比例参与遗传以维持多样性,但必须有 novelty budget、parent-use 标记和降权策略;不能让低质量 novelty 淹没 quality lineage。
- novelty accepted 不等同于 quality progress。若连续 run / generation 只有 novelty accepted 而没有 quality accepted,Agent 必须触发或修复 quality-stagnation rescue,例如 targeted crossover、strategy-module proposal 或结构化参数搜索;不得把 novelty-only 当作扩大长 run 的充分理由。
- 提高吞吐时优先使用 bounded concurrency、cache、duplicate prefilter、cheap validation、operator yield scheduling 和 accepted-candidate quota;不得通过放宽 validation / Reference gate 换速度。
- 在比赛冲刺阶段,structured / LLM strategy-module / Candidate Audit 默认可使用
--workers并行评估;并行度必须显式记录到 artifact,不得省略 post-run gate。 - 修改代码前必须说明目标、涉及文件、预期验收方式。
- 修改文档前必须说明它服务哪个 Source of Truth 文件或 roadmap 阶段。
- 如果 Source of Truth 文档缺失或互相冲突,默认优先补文档和验收计划,而不是扩展代码功能。
9. Evaluation Gates
任何候选进入提交讨论前,必须具备:
- stable
code_hash - local full matrix
- profile matrix
- public anchor
large_seed301风险说明 - runtime 信息
- Reference Portfolio comparison
- validation status
- distill provenance 或 source provenance
large_seed301 是 public anchor risk gate,不是唯一排名标准。小幅退化可以保留为候选,但必须明确风险;明显退化默认不提交,除非 profile matrix 给出强证据。
10. Reference Portfolio
Reference 分三类:
anchor reference:校准线上/线下一致性,如large_seed301和已观测 feedback。quality reference:整体表现较强,用作质量下界。profile reference:在 scarce、low willingness、bundle-heavy、conflict-dense 等局部 profile 上强。
Reference Portfolio 可新增、降权、剔除;每次变更必须记录理由。
当前默认路径:
artifacts/reference_portfolio.jsondocs/audits/phase4-reference-feedback-validation-20260606.md
后续 Agent 不得只凭候选名称或历史印象比较 Reference;必须引用 code_hash、local_matrix_path 和变更理由。
11. Evolution Run Gates
恢复或启动 evolution 前必须确认:
- candidate audit 可运行。
- distill selection report 可生成。
- Reference Portfolio 已建立。
- Online feedback 历史已尽量补录。
- fallback / parser / archive insertion / lineage 日志可审计。
- operator attempt、evaluated candidate、quality Individual、novelty Individual 必须分层记录:attempt 可以失败;quality Individual 必须通过 static / sandbox / duplicate / Reference gate;novelty Individual 必须通过 static / sandbox / duplicate / novelty gate。
- 每个短 run 必须统计 attempt_count、fallback_rate、invalid_rate、duplicate_rate、reference_gate_reject_rate、quality_accept_rate、novelty_accept_rate、LLM latency 和 eval latency。
- 失败率、重复率或 fallback-only generation 超出预算时必须降权对应 operator 或 hard-stop,而不是继续堆 run。
- active training pool 没有未授权的
adv_*.txt污染。 - Phase 6 只允许进入短 Evolution 修复试跑;长 Evolution Run 仍需 Gate 7。
- Gate 7 通过只能放行受控扩大 run;不得把短跑 gate 误解为允许无限制长跑或直接 Online Submission。
- 历史
insights.jsonl已发现 fallback 污染,不得无过滤作为 prompt guidance 复用。 - Reference comparison prompt guardrail 不等于 hard gate;若 Candidate Audit 显示 child 明显弱于 Reference Portfolio,不得作为 Distill / Submission 候选。
Adversarial Instance 只能作为 diagnostic / stress-test;不得静默进入训练池。
12. Distill Gates
Distill 是 AutoSolver 内部 Distiller 的行为,不是外部 AI 编码工具的手工选择行为。Agent 可以实现、修复、审查 Distiller,但只有实际调用 AutoSolver Distiller 并留下 selection report 的结果,才能称为 AutoSolver Distilled Solver。
Distill 必须输出 selection report,至少包含:
- candidate list
- source / island / candidate_id
- code_hash
- local anchor
- local full score
- grouped/profile scores
- selected reason
LLM fusion 只能在 fused Solver 不明显劣于 best candidate 的 public anchor 和 matrix 时接受。
13. Online Submission Gates
每次 Online Submission 前必须写 prediction:
- code_hash
- local matrix
- Reference comparison
- expected strengths / weaknesses
- why this submission is worth one scarce attempt
提交后必须写 postmortem:
- online result
- local-online deviation
- hypothesis confirmed / rejected
- Reference Portfolio 更新
- Scoring Model / Evolution Fitness 是否需要调整
历史提交结果应尽快补录到同一反馈数据库。
当前 Online Feedback DB 默认路径:
artifacts/online_feedback_db.jsonl
缺少 solver.py 的 historical feedback 必须标记为 code_hash_status = unknown,不得推断或借用其他候选的 code_hash。
14. Documentation Rules
- 新文档必须中文撰写,关键术语保留英文。
- 新文档不得直接复制旧文档结论,必须重新审核后写入。
- 每个设计机制必须写清:设计意图、状态标签、已知失败模式、目标实现形态、AI 行为约束、验收标准。
- ADR 仅记录新体系下已审核的新决策;旧 ADR 已归档,不自动继承。
15. Commands And Environment
项目命令默认使用 uv:
UV_CACHE_DIR=.uv-cache uv run python -m src.main status
UV_CACHE_DIR=.uv-cache uv run python scripts/audit_run_candidates.py --run-id <run_id> --include-distilled
Solver sandbox Python 应来自本地 Python 3.6 环境,具体路径以 config.toml 或命令行参数为准。
