Imported from zongxin1993/OnlyAlpha (
AGENTS.md). Install upstream withnpx skills add zongxin1993/OnlyAlpha. Copyright stays with the author.
OnlyAlpha Agent 工程指南
本文件定义 OnlyAlpha 每个工程任务的强制执行与验收规则。它规定 开始前必须理解什么、如何开发、如何验证、何时停止,不记录任何任务、阶段或版本的完成状态。
0. 首要规则:必须先读 Project Constitution
任何 Codex / Agent 在执行以下工作之前:
architecture analysis
planning
audit
implementation
refactor
bug fix
review
test design
ADR drafting
roadmap interpretation
MUST 首先完整阅读并理解根目录 PROJECT_CONSTITUTION.md。
然后按以下顺序获取上下文:
1. PROJECT_CONSTITUTION.md
2. 相关 Architecture / public Contracts
3. 相关 Accepted ADRs
4. Roadmap / 当前任务上下文
5. 本 AGENTS.md 的工程执行规则
6. 当前源码 + 当前测试 + 当前可执行行为
其中:
PROJECT_CONSTITUTION.md
→ L0 最高规范性 Authority:OnlyAlpha 为什么存在、最终是什么、不可牺牲原则和永久边界
Architecture / Contract
→ L1:如何实现 Constitution 定义的目标
Accepted ADR
→ L2:局部设计决定,只能在 L0/L1 之下生效
Roadmap / Work Program
→ L3:建设顺序与任务分解,不拥有改变产品目标的 Authority
Task Contract / Prompt
→ L4:当前授权实现范围
当前源码 + 当前测试
→ observational truth:当前工程实际上实现了什么
0.1 Constitution 不可由普通工程任务修改
Codex / Agent / 普通实现任务 MUST NOT:
- 修改、删除或重写
PROJECT_CONSTITUTION.md; - 修改其 pinned fingerprint;
- 通过 ADR supersede Constitution;
- 因实现困难降低 Constitution 中的 Required Principle;
- 把“当前还没实现”解释成“长期目标已取消”;
- 把局部 sequencing 优化解释成产品 scope 缩减;
- 通过 Prompt、测试或实现重新定义 OnlyAlpha 的产品身份。
若任务、现有 ADR、Roadmap、代码或测试与 Constitution 冲突:
STOP IMPLEMENTATION
REPORT: PLAN_CONFLICT
Agent 必须说明冲突来自哪里、违反哪一条 Constitution、需要什么 owner-level 决策;不得自行通过修改 Constitution 解决冲突。
0.2 Normative Truth 与 Implementation Truth 必须分离
Constitution / Architecture / Contract
→ OnlyAlpha 应当成为什么
Current source / tests / executable behavior
→ OnlyAlpha 当前已经实现了什么
源码可以证明 Futures、某插件、某节点“当前没有实现”;源码不能证明它“不是长期目标”。
1. OnlyAlpha 永久架构不变量
以下规则必须与 PROJECT_CONSTITUTION.md 一致,并在每个相关任务中保持:
- OnlyAlpha 是长期运行的个人 Stateful Quant System,而不是以 Python callback 为正式产品边界的量化脚本框架;
- Trading Kernel 只拥有不会因具体市场规则变化而变化的 canonical trading semantics;
- Core 必须 Market-Agnostic,不依赖具体 Provider / Broker / Exchange / Market 实现;
- 任何会随市场、交易所、Broker、Provider、监管、协议版本变化的内容,必须停留在 Plugin / Adapter / Gateway 边界;
onlyalpha.domain是跨市场 canonical domain language;- RESEARCH / BACKTEST / SIM / LIVE 是正式 Runtime vocabulary;
- Backtest / SIM / LIVE 共享 Trading Kernel 与同一交易语义;
- Strategy fingerprint 是策略身份;Freeze 后产生 immutable Strategy Revision;
- 同一 Strategy Revision 进入 Backtest / SIM / LIVE,Runtime 不重新定义策略语义;
- 一个正式语义事实只有一个 canonical identity 和一个 Authority,不得维护平行真相;
- 外部 venue 是 execution fact authority;OnlyAlpha 拥有 intent、policy、strategy identity、promotion、reconciliation 与本地可恢复状态 authority;
- 相同 Kernel version + canonical configuration + Strategy Revision + initial state + ordered facts 必须得到相同 state transition 和 output;
- 所有影响结果的时间、随机性、模型版本、外部结果、人或 Agent 决策必须成为显式输入、版本配置或正式事实;
- 市场与执行事实 append-only;修正产生 revision,不静默重写历史事实;
- Research / Backtest 正式 Evidence 必须绑定 immutable inputs,不能依赖 mutable database query 作为唯一数据定义;
- 所有正式链路必须可 Trace:能够从 Fill 追溯到 Order Intent、Risk/Portfolio/Strategy Decision、Strategy Revision 和输入事实,也能够正向追踪影响;
- Crash / Restart 是正常生命周期;关键 Truth 不能只存在内存,必须能从 durable facts + reconciliation 恢复到唯一状态;
- UNKNOWN submit outcome 是一等状态,禁止 blind retry 创建第二订单身份;
- LIVE 对新的 risk-increasing execution fail closed,但必须继续观察、成交、撤单、持久化、reconciliation、recovery 和风险降低链路;
- Web 只负责 Display、Input/Management、Command Submission;UI/Web 永远不是 Trading Authority;
- 正常 Human 操作通过 Web → versioned API → OnlyAlpha,不通过直接 Python Core 调用或数据库修改;
- Agent 只能通过正式 API 使用 OnlyAlpha,职责限定于 Factor 实现/挖掘/管理、Research、Backtest、SIM 与 Evidence 分析;
- Agent 永远不拥有 LIVE Authority;LIVE activation、LIVE strategy change、material LIVE risk authorization 必须人工明确操作;
- Infrastructure 负责 node identity、deployment boundaries、interfaces、compatibility、health、upgrade、rollback、persistence topology、observability、failure domains 与 lifecycle;具体 Docker/Kubernetes/DB/Web 技术不是 Constitution;
- 跨节点通信使用明确、版本化 Contract;Database 不是默认 integration API;
- correctness 测试使用 deterministic barrier、event、fake clock 或 fault injection,禁止用
sleep()证明正确性; - Convenience、性能和交付速度不能作为越过 Authority / Boundary / Determinism / Recoverability 的理由。
1.1 代码归属第一判断
新增能力先问:
Will this change because an external market/provider rule changes?
如果 YES,默认属于 Plugin / Adapter / Gateway。
如果 NO,继续判断它是否属于 universal canonical trading semantics;只有属于时才考虑进入 Core。
Human interaction → Web。
Factor discovery / research automation → Agent / Research。
Node / deployment / compatibility / lifecycle → Infrastructure。
若新市场能力无法表达,必须先判断是 provider-specific difference,还是 Core 确实缺失 universal concept。只有后者才构成 Core 演进理由。
2. Authority 模型
OnlyAlpha 固定使用以下工程 Authority 分工:
PROJECT_CONSTITUTION.md
→ 最高规范性 Authority,不可由普通任务 supersede
Architecture / public Contract
→ 由 Constitution 派生的长期结构与接口约束
Accepted ADR
→ 局部设计决策;与上层冲突时无效
当前源码 + 当前测试
→ 当前工程实际实现事实
AGENTS.md
→ 每个任务如何执行、验证、何时停止
quality-policy.toml
→ 持续 CI 与 Major Milestone Phase Gate 的机器检查集合
scripts/test_suite.py
→ canonical lane 的实际执行定义
scripts/verify.py
→ 当前工作区修改的 Impact-Aware 验证选择器
pyproject.toml
→ Python 测试、静态检查与包配置的机器事实
不得创建第二份任务验收 Authority、工程进度 Authority、质量结论 Authority 或认证状态 Authority。
源码与长期设计冲突时不得静默选择一边。首先检查 Constitution:
- 若源码违反 Constitution/冻结设计,修实现;
- 若低层文档违反 Constitution,修低层文档;
- 若任务本身要求违反 Constitution,
PLAN_CONFLICT并停止; - 普通任务无权通过“设计变化”修改 Constitution。
3. 每个正式任务的最小 Task Contract
开始实现前,当前开发/Codex 上下文必须明确:
Goal
Modification Scope
Expected Impact Scope
Required Behavior
Acceptance Tests
Out of Scope
Stop Condition
Constitution Impact
其中 Constitution Impact 必须回答:
Does this task conflict with, weaken, reinterpret, or require changing PROJECT_CONSTITUTION.md?
合法普通任务答案必须是:
NO
若为 YES 或无法确定,停止实现并报告 PLAN_CONFLICT。
Task Contract 只存在于当前任务上下文,不提交仓库,不生成模板实例、完成报告、验收报告或状态文件。
3.1 Required Behavior
Required Behavior 在任务开始时冻结。实现困难不能反向降低 Required Behavior。
开发中发现现有设计不足时,只能在 Constitution 允许范围内显式更新对应 Architecture / Contract / proposed ADR,并同步实现和测试。
Codex 可以起草 PROPOSED ADR,但不得把与 Constitution 冲突的 ADR 当作有效决策,也不得仅因自身实现选择把产品 scope 缩小。
3.2 Acceptance Tests
Acceptance Tests 可以因为真实发现而增强或纠正,例如:
- 新发现的真实边界条件;
- 实际依赖使 Impact Scope 扩大;
- 原测试不足以证明 Required Behavior;
- 原测试本身被证明错误;
- 需要新增 determinism、recovery、compatibility 或 migration 证明。
不得因为实现无法满足要求而删除测试、弱化断言、放宽语义、增加无意义 retry/sleep、skip/xfail 当前真实失败或吞掉异常。
3.3 Web / Product Vertical Slice Tasks
Web 产品化任务除通用 Task Contract 外,必须遵守 ADR 0135 与 docs/web-product-development-mode.md。
在开始实现之前,当前上下文必须额外明确:
Product Slice
Reference Interaction
User Goal
Primary User Flow
Visible States
Existing Product API
Missing Product API / Authority Capability
E2E Acceptance
默认开发链必须是:
用户可见工作流
→ UI / Product Slice
→ 当前 Product API / Authority Gap Analysis
→ 优先复用已有能力
→ 只补当前 Slice 所需的最小完整 canonical capability
→ formal Product API
→ Web integration
→ Browser E2E
→ dogfooding / review
→ closure
禁止把 Web 产品化重新退化成“先横向完成大量后端模块,最后再给它们加页面”。若当前 Product Slice 不需要某个后端能力, 默认不为假设未来需求提前增加 API、Authority、abstraction、configuration 或 persistence surface。
上述约束不阻止独立必要的 correctness、安全、数据完整性、recovery/reconciliation、quality infrastructure 或 owner 明确批准的 非 Web roadmap 工作。
TradingView 只作为 chart-centric interaction reference。它不拥有 OnlyAlpha 产品语义、Domain、API 或 Authority,也不自动授权 替换现有 renderer。ADR 0094 的 renderer boundary 在被新的 Accepted ADR 显式 supersede 之前保持有效。
Web 仍然只能是 Control + Presentation。若页面需求不能通过当前正式 API 表达,必须先判断是 PRESENTATION_GAP、QUERY_GAP、 COMMAND_GAP、DOMAIN_GAP、AUTHORITY_GAP 还是 INFRASTRUCTURE_GAP,再修改 canonical boundary;禁止通过 React 直连数据库、 Store、internal Core 或浏览器重算 Research/Trading truth 规避缺口。
Web Slice 的默认 closure 不是“endpoint 已存在”或“component 能 render”,而是主用户路径通过 formal Product API 端到端可用, 并有与风险相匹配的 Browser E2E / targeted evidence。
3.4 Simplicity / Accidental Complexity Control
OnlyAlpha 使用 bounded Simplicity Review 消除 accidental complexity;它不删除 Constitution、Architecture / Contract、Accepted ADR、冻结 Required Behavior、correctness、security、reproducibility、data integrity、recovery、traceability、observability 或 required tests 所要求的 essential complexity。
对涉及 executable code 的非平凡实现、重构或 Bug 修复,Agent 在理解真实调用链和边界之后,按以下顺序选择实现:
1. 当前 Required Behavior 是否真的需要新增实现?若不需要,不新增。
2. 当前代码库是否已经有可复用实现?优先复用。
3. Python / JavaScript / 平台标准能力是否已经覆盖?优先 stdlib / native。
4. 已安装依赖是否已经可靠覆盖?优先复用,不为同类能力新增 dependency。
5. 只有以上都不成立时,新增满足当前 Required Behavior 的最小正确实现。
默认禁止仅为了假设未来需求新增 abstraction、interface、factory、adapter、wrapper、configuration switch、extension point 或 dependency。第二个真实实现、明确 Contract、已接受 Architecture/ADR 或当前 Required Behavior 可以构成新增抽象的证据。
“最小”指最小必要复杂度,不是 code golf。不得用更短代码换取更差的可读性、边界正确性、类型安全、错误语义或确定性。
若当前 Codex 环境已安装 Ponytail,默认使用 Full semantics;实现阶段可调用 @ponytail,在代码修改完成并通过初始 Baseline Validation 后调用 @ponytail-review。若 Ponytail 不可用,Agent 仍必须直接执行同等 bounded Simplicity Review;第三方插件可用性不得改变 Required Behavior、验证范围或 Stop Condition。
Simplicity Review 只检查当前 Modification Scope + 真实 Impact Scope 内的 diff,重点寻找:
- 可删除的 dead/speculative code;
- 已存在项目实现却重复实现的逻辑;
- stdlib/native 已覆盖的手写实现;
- 为单一实现/单一调用方提前建立的抽象层;
- 没有当前消费者的配置、扩展点和 dependency;
- 保持相同语义时可以明显缩小的实现。
每个 finding 必须在当前开发上下文中二选一:
APPLY
→ 简化实现,并重新执行受影响的 targeted tests / static checks
REJECT_AS_ESSENTIAL
→ 说明它由哪条 Required Behavior / Architecture / correctness / safety / reproducibility / recovery 约束要求保留
不得把 Ponytail finding、@ponytail-audit 输出、debt ledger、PASS 结果或 review summary 提交成第二份仓库质量/进度 Authority。@ponytail-audit 也不得把普通任务扩展成全仓 cleanup;只有当前 Task Contract 或 Major Milestone contract 明确把 repository-wide complexity audit 纳入 Scope 时才能执行。
Ponytail 不拥有任务验收 Authority,不替代 Baseline Validation、risk-specific evidence、bounded Independent Review、Constitution consistency 或现有 CI/quality gates。
4. 验收模型:Risk-Tiered + Impact-Aware
验证范围由 真实 Impact Scope 决定,不由任务编号决定。
4.1 普通任务 Baseline Validation
涉及代码的普通任务默认至少执行:
- 与修改直接相关的 targeted tests;
- affected Ruff
check; - affected Ruff
format --check; - 涉及 Python 类型或 API 时,对 affected scope 执行 mypy;
- 触碰已有 subsystem 时,执行最近的 affected canonical lane;
- 若触碰 architecture / contract / boundary,检查 Constitution consistency。
随后只按真实 Impact Scope 增量扩展。不得因为“更保险”自动执行全仓测试、完整 Layered Quality、CodeQL、全 build matrix 或全量 coverage。
纯文档修改不需要无关 Python 测试;但修改 ADR / Contract / governance 时必须检查其直接依赖和架构一致性。
4.2 高风险任务判定
只要修改失败可能破坏以下任一性质,就按高风险任务处理:
- 资金安全;
- execution correctness;
- 持久事实完整性;
- recovery / reconciliation 能力;
- 唯一 Authority;
- determinism / identity;
- public compatibility;
- security boundary;
- Constitution / architecture governance;
- quality infrastructure。
典型高风险范围包括:
Authority / ownership
Trading state machine
Persistence / Schema / Migration
Checkpoint / Recovery / Reconciliation
Order / Execution
Broker
Risk
LIVE safety
Public Core Contract / SPI
Wire / serialization contract
Compatibility boundary
Security boundary
Governance / quality infrastructure
4.3 高风险任务额外要求
高风险任务在 Baseline Validation 基础上增加真实相关专项验证,并进行一次 bounded Independent Review。
Independent Review 重点检查:
- Constitution 是否被违反或弱化;
- Authority 是否唯一;
- 状态机是否存在非法状态;
- fail-closed 是否成立;
- recovery / retry 是否确定;
- public contract 是否被静默改变;
- 跨模块边界是否被穿透;
- 是否存在绕过 API、测试或正式入口的隐式路径。
Review 范围严格限制为:
Modification Scope
+ 真实 Impact Scope
+ 直接相关 Constitution / architecture invariants
禁止借 Independent Review 重新启动全仓审计。
5. Scope、Severity 与 Hard Stop
Severity:
Critical → 阻塞
High → 阻塞
Medium → 默认不阻塞
Low → 不阻塞
任何直接违反 Constitution 的问题至少是 High;涉及资金安全、Authority 冲突或可能产生不可恢复真实交易状态时按实际风险提升到 Critical。
当前任务只修改:
Modification Scope
+ 因真实依赖关系证明必须扩展的 Impact Scope
发现范围外真实问题默认不修,除非它阻止 Required Behavior 或证明原 Impact Scope 判断不完整。Impact Scope 只扩到最近稳定工程边界,禁止无限审计。
普通任务满足:
Required Behavior 已实现
+ Acceptance Tests PASS
+ Baseline Validation PASS
+ 真实 Impact Scope 所需验证 PASS
+ bounded Simplicity Review 已完成(涉及非平凡 executable code 时)
+ Constitution consistency PASS
+ 当前范围 Critical = 0
+ 当前范围 High = 0
= STOP
高风险任务额外要求:
bounded Independent Review 完成
= STOP
达到 Stop Condition 后,不得因为 speculative risk、Medium/Low、无关技术债、“还能优化”或“还能重构”继续扩大当前任务。
6. Determinism、Tests 与 Evidence
默认 Task Acceptance 测试必须 deterministic、hermetic、offline-first。
普通验收不得隐式依赖公网、实时市场数据、真实时间推进、第三方账户状态或不可控执行顺序。
优先使用 fixture、recorded deterministic payload、fake clock、contract fake、local ephemeral DB、deterministic barrier 与 controlled fault injection。
禁止通过以下方式“跑绿”:
- retry-until-green;
- 增加
sleep(); - 无依据扩大 timeout;
- 吞掉异常;
- 依赖随机执行顺序;
- 删除有效断言;
- 降低断言强度;
- 无依据扩大 tolerance;
- 删除有效边界测试;
- skip/xfail 当前真实失败;
- 修改测试去适配违反 Contract/Constitution 的实现。
确认真实 Bug 原则上必须增加最小 Regression Test:
复现旧错误
→ 修复根因
→ 在最近稳定边界证明不会复发
真实 Binance、QMT、CTP、数据库部署、Docker 等环境只有在它们本身是 Required Behavior 的不可替代证明时才是当前任务强制验收项。环境不可用不是 PASS;Mock/fake 不得冒充必须由真实环境证明的集成行为。
历史已有失败不自动阻塞当前任务,但缺少充分证据也不能等价为 PASS。
Coverage 是诊断与专项验证工具,不是普通任务默认 Gate,不得为了固定百分比制造低价值测试。
本地正式验证必须通过 deploy/run-tests.sh 进入唯一 canonical Compose test profile;宿主 .venv 只允许用于快速 inner-loop
定位,其结果不得替代正式验收。Runner 必须构建当前工作树、使用隔离 namespace、保留 test-results/,并在退出时清理容器、
网络和测试卷。明确依赖公网、Windows、硬件或真实账户的 external lane 在对应受控目标环境执行,不得伪装成容器 PASS。
本地 PostgreSQL 测试必须通过唯一 canonical deploy/docker-compose.dev.yml 的 test profile 执行;
不得在宿主机手工启动 PostgreSQL 或为测试引入第二套 PostgreSQL DSN 变量。Compose
容器只使用正式连接配置 ONLYALPHA_POSTGRES_DSN,并必须保留 _test 数据库后缀安全校验。
当前开发阶段,deploy/docker-compose.dev.yml 是 OnlyAlpha 唯一授权的 Docker Compose 拓扑。
新增第二个 OnlyAlpha-owned Compose topology 必须先取得明确的 owner approval,并同步更新本治理规则与架构测试。
7. Compatibility、Persistence、Build 与 Security
任何公共 Contract 或持久格式兼容性变化都属于高风险任务,包括:
- public Python API;
- Provider / Broker / DataSource SPI;
- versioned external API;
- wire protocol;
- persistent schema;
- checkpoint / recovery format;
- event serialization;
- plugin contract;
- public CLI behavior。
必须判断 backward/forward compatibility、migration 与 affected consumers,并按需执行 contract、consumer、migration tests 与 bounded Independent Review。
Breaking change 可以存在,但必须是 Constitution 允许范围内的明确设计决定,不能因为实现方便静默发生。
活跃开发阶段不默认承诺 backward compatibility。只有 repository owner 明确要求,或 Contract 明确标记为 compatibility-frozen / externally supported 时才必须保留;默认不得增加 compatibility shim、version-family duplication 或 legacy loader。Breaking-change detection、分类及仓内 consumer 原子迁移仍然必须执行。
数据库 migration 必须明确 precondition、deterministic transformation、failure semantics、transaction/atomicity、compatibility window、restart/retry semantics 和 data integrity。能安全回滚时提供 rollback;不能安全回滚时使用 fail-closed + backup/snapshot/forward-fix。
修改 package metadata、dependencies、entry points、public exports、plugin discovery、frontend build inputs、Docker image contents 或 release/build scripts 时,执行对应 build/package 验证。
认证授权、secret handling、外部输入、网络协议、SQL/persistence、命令执行、文件系统、Web security、Broker/LIVE 外部接口等修改按真实风险增加专项 security 验证。
8. Quality / Governance Infrastructure Protection
以下属于质量或治理基础设施:
PROJECT_CONSTITUTION.md
docs/governance/*
AGENTS.md
quality-policy.toml
scripts/verify.py
scripts/test_suite.py
scripts/check_constitution.py
.github/workflows/* governance/quality rules
architecture rules
lint / mypy / test discovery configuration
业务/功能任务不得为了让当前实现通过而修改质量规则、ignore、allowlist、threshold、test discovery、gate selection 或 workflow condition。
PROJECT_CONSTITUTION.md 与其 fingerprint 对普通任务绝对只读。
其他质量/治理基础设施只有在 Task Contract 本身明确以该基础设施为修改目标,或有工程证据证明现有规则本身错误且修改不违反 Constitution 时才能修改,并按高风险任务验收。
verify.py 是当前工作区修改的 Impact-Aware Verification Selector,不是任务状态机、长期认证系统或产品规划 Authority。
9. 文档与仓库卫生
9.1 Repository Semantic Identity
永久仓库 identifier 必须描述稳定的 domain semantics、responsibility、observable behavior、schema change 或 invariant, 不得描述它来自哪个 roadmap phase、implementation task、closure round、milestone、temporary migration stage、PR 或 issue。
本规则适用于 src/、packages/、plugs/、database/、tests/、test-data/、scripts/、examples/、
contracts/ 与 deployment / CI 资产中的路径、模块、类型、函数、fixture、helper、test name、parametrize ID、
golden/case ID、migration suffix、SQL object、error code、runtime resource 和长期配置名。
例如:
test_b3_registry.py → test_factor_registry.py
test_p9_case_2() → test_unknown_order_is_reconciled_without_duplicate_submission()
0010_p9_0_authority_hardening.sql → 0010_strategy_revision_authority.sql
a0_binance_golden → binance_spot_price_filter
Migration 的数字顺序前缀、正式 schema/API version、稳定 domain 编号以及 exact replay / immutable identity 所要求的 既有序列化身份不是开发过程标签。此类例外必须由当前 Architecture / Contract / Accepted ADR 明确要求,并在机械 Gate 中 使用精确 token 或精确 path 记录原因;禁止目录级、glob 或模糊 allowlist。
过程文档(docs/tasks/、docs/plans/、docs/audits/、docs/handoffs/、Roadmap、Prompt)可以保留过程身份,
因为其职责就是记录开发历史。Git 记录何时发生;源码、测试、迁移与 fixture 描述当前产品事实。
仓库只保存系统长期需要的信息:
- Project Constitution;
- 源码;
- 测试;
- ADR;
- Architecture;
- Contract / protocol;
- 必要用户、开发、部署说明;
- Roadmap / execution plan 中的未来建设顺序和依赖关系。
仓库不得把以下内容作为当前 Authority:
- 每步完成状态;
- 工程进度状态文件;
- 质量/审计/验收/closure 报告;
- CI PASS 快照;
- verification manifest/history;
- Final-SHA / Exact-SHA 工程认证记录;
- task implementation/validation/completion summary;
- 自动生成的下一步授权状态;
- 历史 Prompt。
Roadmap 只能描述建设地图和依赖关系,不得通过 completed/current/pending/ready/verified/progress percentage 取代当前源码事实。
历史 Prompt 只能解释历史任务,不拥有产品目标、当前状态或下一步工作 Authority。
10. CI、Phase Gate 与默认执行顺序
GitHub CI 是持续质量探针,不是每一步任务的默认完成 Authority。
CI pending → 不阻塞普通任务
CI PASS → 附加质量信号
CI FAIL → 判断是否为当前修改导致的真实 regression
只有已确认当前修改引入真实 regression,或某项验证只能在 CI/特殊环境完成且是 Required Behavior 不可替代证明时,CI 才成为当前任务阻塞条件。
完整 Phase Gate 只在 Major Milestone 真正结束时执行一次,并按 quality-policy.toml 和 milestone contract 执行适用的 repository-wide static、canonical regression、architecture、recovery、conformance、persistence、build、E2E 与 security 检查。
面对任何工程任务,默认顺序:
READ PROJECT_CONSTITUTION.md
→ READ relevant Architecture / Contracts / Accepted ADRs
→ READ AGENTS.md
→ inspect current source/tests
→ Constitution Impact check
→ if conflict: PLAN_CONFLICT + STOP
→ establish Task Contract
→ freeze Required Behavior
→ implement smallest correct solution under Simplicity Discipline
→ correct Impact Scope from real dependencies
→ targeted tests + Baseline Validation
→ bounded Simplicity Review when non-trivial executable code changed
→ if simplicity changes code: rerun affected validation
→ add risk-specific evidence
→ bounded Independent Review when high-risk
→ Constitution consistency check
→ Critical = 0 / High = 0
→ STOP
核心原则:
Constitution 定义 OnlyAlpha 必须成为什么;代码工程定义 OnlyAlpha 当前做到了什么。实现必须充分,但永远无权通过局部实现重新定义项目愿景。 每次读到这里,要求在当前在任务的最后,输出现在的北京时间
Repository placement and naming
Any concrete implementation loaded through the OnlyAlpha Plugin SPI MUST live under
plugs/onlyalpha-plugin-<NAME>/. Any independently buildable, versioned, deployable
non-plugin component MUST live under packages/onlyalpha-<FUNC>-<NAME>/. Stable,
market-agnostic canonical semantics and the Plugin SPI belong under src/onlyalpha/.
Do not add apps/<...> or category-first wrappers such as packages/provider,
packages/market, packages/api, packages/protocol, packages/factor,
packages/indicator, packages/target, or packages/fake. Do not place a non-plugin
component under plugs/, and do not place a concrete plugin under packages/.
If a new architecture requires a different boundary, record an explicit architecture
decision before adding the path. Web and HTTP transport are components, not plugins;
the OpenAPI Product Contract remains a first-class contract under contracts/.
Quantitative asset placement
Classify each new quantitative capability under ADR 0110:
- Generic mathematics without financial context is an Operator.
- Stable financial meaning without a predictive Target hypothesis is an Indicator.
- A testable predictive or explanatory hypothesis is a Factor.
- Composition of admitted Features/Factor into eligibility, selection, entry or exit decisions is a Strategy.
Operator/Indicator are public reusable capabilities. Production Factor/Strategy assets are private; the main repository's examples are
non-production DB import seeds under examples/private-assets/factor/ and examples/private-assets/strategy/. The Agent primarily
creates/searches Factor/Strategy. Missing reusable Operator/Indicator capability must be
proposed and admitted separately, never hidden inside a Factor or Strategy.
Under ADR 0129, ADR 0131 and ADR 0132, production Private Factor/Strategy authoring is PostgreSQL-backed and uses mutable Drafts plus immutable Revisions. Private Factor V1 is one canonical UTF-8 Python source unit bound to one exact stable Factor API Contract; Private Strategy is a canonical structured Strategy Definition. Git repositories, source paths, editable installs, wheels and packages are optional interoperability/provenance paths, not production authoring Authority. The database source is never direct execution permission; exact API/adapter, validation, Provider, Catalog and Runtime Generation boundaries remain mandatory.
ADR 0111's source/distribution loading remains valid for optional import/export and executable materialization. Do not reintroduce mandatory per-asset Git/package authoring workflows without a new ADR; private Factor/Strategy planning must follow ADR 0129–0132.
Operator/Indicator distributions and snapshot-backed Factor providers use onlyalpha.quant_assets under ADR 0112. Operator/Indicator/Factor
execute only through onlyalpha.calculations; Strategy remains authoring data. Any content change requires a new provider version, and any semantic
change additionally requires a new Calculation or Strategy-asset semantic version. Hot plug switches an immutable catalog generation only
for new work; never reload modules in place or rebind an active Run/StrategyRevision.
Public example / private asset contract parity
When a public Core change affects a Factor/Strategy authoring, discovery, execution, Research, Evidence or Freeze contract, the implementer must inspect the corresponding DB-native seed, update it in the same public change when behavior changes, and run the seed import/native execution conformance lane.
When a private asset needs a new Core capability, it must first be expressible through a public OnlyAlpha contract, demonstrated by the
corresponding public example, and consumed through that same contract. Hidden private-only Core integration paths are forbidden and
must fail closed as EXAMPLE_CONTRACT_COVERAGE_REQUIRED until public contract/example coverage exists.
Test placement standard(测试目录与测试数据标准)
tests/ 采用单一分类法:按被测模块区域分目录,测试目的用 pytest marker 与区域子目录表达,禁止用平行目录表达目的。
测试代码归属
src/onlyalpha/<area>/的测试默认落tests/<area>/;同一模块的测试 MUST NOT 分散到多个历史平行目录(如tests/unit/<area>、tests/<area>_conformance与tests/<area>并存)。合并时保留 git 历史(git mv)。scripts/工具自身的测试落tests/tools/;packages/<component>/与plugs/<plugin>/的测试只落各自组件目录内的tests/,根tests/不接受组件专属测试。- 跨切验证区域(architecture、conformance、contracts、integration、scenario、performance、property、formal、certification、examples、deploy、cli 等)各自承载一种验证视角,不得为同一模块再造目的重复目录;新增跨切区域必须在节内登记。
- unit / contract / recovery / scenario 等目的区分使用
pyproject.toml已注册的 pytest markers,不用目录层级复制。 - 每个测试包目录 MUST 含
__init__.py;共享测试辅助代码统一放tests/support/(可按职责建子包);被 subprocess 以模块方式拉起的测试入口脚本统一放tests/runtime_support/。tests 根目录 MUST NOT 存放非测试模块的散文件。 - 空壳目录(无受跟踪测试文件的目录,含仅存
__pycache__)属于遗留垃圾,发现即删除;tests/内 MUST NOT 出现.DS_Store等系统文件被跟踪。
测试数据归属
测试数据只允许两个落点:
tests/<area>/fixtures/:唯一消费者是某区域的场景绑定数据,与该区域 co-locate,随场景在同一 commit 原子更新。test-data/:跨多个区域消费的数据、golden 基线(如 recovery baselines)、共享参考向量包与生成型工件(如test-durations.json)。子目录按 owning domain 命名并保留README级可发现性,MUST NOT 倾倒为无结构散件。
禁止新增第三类测试数据目录(如 tests/fixtures/、tests/reference_data/ 式的根内异构数据树)。与代码原子演进的 golden 数据必须由再生脚本 + 一致性守卫(fail closed)绑定,路径变更 MUST 同步更新守卫与再生脚本。生成物与测试产物缓存(test-results/、.test-cache/)保持 gitignored,不入 test-data/。
存量迁移按上述两规则执行;每次触碰某数据目录时按本标准收敛,不做无验证范围的全仓一次性搬迁。
External engineering evidence and upstream failure research
External projects, upstream documentation, issues, pull requests and incident reports are engineering evidence only; they never supersede PROJECT_CONSTITUTION.md, OnlyAlpha Architecture / Contracts / Accepted ADRs, current Task Contract or current implementation truth.
Before freezing architecture or a repository-aware plan, Codex / Agent MUST use first-principles reasoning first and then consult relevant curated material under docs/engineering/ when mature external references exist for the affected domain.
For high-risk work, a new domain, provider/venue semantics, unresolved correctness defects, or a new certification boundary, planning MUST inspect relevant upstream issue/bug history when such evidence is available. The research is bounded by the real Impact Scope and MUST NOT become an open-ended ecosystem audit.
The required transformation is:
external design / issue / incident evidence
→ applicable generalized lesson or failure pattern
→ OnlyAlpha invariant / failure semantic
→ regression, differential or fault-injection test
Borrow failure modes and invariants, not patches. A third-party retry, timeout, sleep, mutable status model, evaluator, runtime choice or component graph is never adopted merely because a mature project uses it.
When external evidence materially affects a high-risk plan, the plan should explicitly record:
Relevant references
Applicable design lessons
Applicable historical failure patterns
Rejected reference approaches
OnlyAlpha invariants / tests derived from the evidence
Routine planning SHOULD consult docs/engineering/reference_registry.md and docs/engineering/failure_patterns/ before live upstream research. Live research is required only when risk, novelty, provider change or unresolved defects justify it.
A confirmed applicable external failure SHOULD become an executable OnlyAlpha regression/fault test when deterministic reproduction is practical. Documentation such as "be careful with reconnect" is not an adequate substitute for a test when the failure can be mechanically proved.
The detailed method is defined in docs/engineering/open_source_engineering_evidence.md. These reference documents remain advisory engineering knowledge and MUST NOT become a second architecture, quality, Research, execution or certification Authority.
Formal proof / Authority-bound audit
以下规则补充第 4.3 节的高风险 Independent Review 与第 5 节 Stop Condition。它们适用于会影响正式 Authority、canonical identity、immutable fact/history、projection、Evidence/provenance、state machine、recovery/replay、exact query/certified absence、fail-closed、admission/reuse/suppression/decision witness 的 proof-bearing / Authority-bound 任务。
这类任务在冻结实现前,MUST 从 owning Authority 与 Accepted Contract/ADR 重新推导,而不是只从本次 bug、当前代码分支或实现 checklist 出发。至少完成:
Authority State-Space
→ 枚举真实 Impact Scope 内所有合法 state / outcome / terminal variant
Proof Matrix
→ 对每个 mandatory dimension 明确 Authority、canonical identity、representation、validation、missing semantics
Relation Closure
→ 明确 occurrence / ownership / lineage 的 authoritative relation chain;共享 Result/ref/时间戳不得替代关系证明
Predicate Truth Table
→ 区分 proved match、proved non-match、incomplete proof、unavailable proof
Negative Invariants
→ 明确哪些结论绝不能从缺失、UNKNOWN、共享 identity、cancel/stop/failure 等事实推导
Structural Mutation
→ 至少覆盖 whole-context、owner identity、nested relation、source ref、leaf identity、duplicate owner、wrong family、complete-different identity
若 formal predicate 内部使用 True / False / None,语义 MUST 为:
True → 已完整证明 exact match
False → 已完整证明 exact non-match
None → 缺少完成判断所需的 mandatory proof
因此:
每一个
False都必须有完整的 non-match 证明;missing evidence / malformed relevant proof 不得作为 False,也不得被升级为 certified absence。
对 typed owner / exact predicate,默认按以下三阶段推导:
Applicability
→ record 是否属于该 predicate / owner family?
Completeness
→ 若适用,mandatory proof 是否完整?
Equality
→ proof 完整后 exact identity / context 是否匹配?
对于状态机,必须同时审计 success / failure / cancellation / incomplete / UNKNOWN 等合法状态,不能只覆盖 happy path;optional downstream fact 不得成为 occurrence 存在的隐式前提。
高风险 proof-bearing 任务的 bounded Independent Review MUST 独立重推导相关 Authority state space、relation closure、negative invariants、predicate truth table 和 adversarial mutations,再与实现比较;不得仅复读实现 Prompt 的 checklist。
此类任务在第 5 节原 Stop Condition 之外,适用项还必须在当前任务上下文证明:
Authority State-Space PASS
Entity Completeness PASS
Relation Closure PASS
Predicate Proof Sufficiency PASS
Negative Invariants PASS
Structural Mutation PASS
Recovery / Replay PASS(涉及 durable history / recovery 时)
Independent Re-Derivation PASS
并保持当前范围:
Critical = 0
High = 0
不适用项可以 N/A,但必须说明为什么不在真实 Impact Scope;不得用 N/A 跳过实际 proof obligation。上述 PASS/N/A 只存在于当前 Task / Review 上下文,禁止提交成仓库质量状态、closure 报告或第二份验收 Authority。
详细执行方法见 docs/engineering/formal_proof_audit_method.md。该文档解释如何执行本节规则,不拥有独立的 Task Acceptance Authority。
