Imported from linofcp007/dev-spec-driven (
AGENTS.md). Install upstream withnpx skills add linofcp007/dev-spec-driven. Copyright stays with the author.
dev-spec-driven — agent instructions (tool-agnostic)
This file is the portable version of the dev-spec-driven workflow. Any agent tool that reads an
instructions file — Codex CLI, Gemini CLI, Cursor, Windsurf, Copilot, Claude, Zed, Cline, … —
can follow it. The full reference lives in skills/dev-spec-driven/SKILL.md and references/.
Language: detect the user's language and respond in it (English, Português, Español), including the prose inside generated artifacts. Pass
--lang en|pt|estodev-spec init(sets the project default) anddev-spec create(per feature; inherits the project default) so the scaffolds, steering and tool messages come out already localized — you only fill the placeholders. Keep structural tokens stable (AC IDs likeUS-1.AC-1, task markers_Requirements:_, track namescore/+tdd/+saas/+ai). EARS keywords work in all three:SHALL/DEVE/DEBE,WHEN/QUANDO/CUANDO, etc.
What this is
Spec-driven development that scales rigor to the feature. You classify each feature into composable
tracks, then run an approval-gated pipeline producing traceable artifacts in .specs/.
| Track | Adds | Turn on when |
|---|---|---|
| core (always) | EARS requirements → design → tasks → execute | every Spec-mode feature |
| +tdd | test plan + failing-tests-first + red→green→refactor | correctness matters / hard to undo |
| +saas | performance/scale/multi-tenancy/observability/cost + load test | multi-tenant, hot path, prod scale |
| +ai | eval-driven dev, prompts-as-code, token economics, safety | quality depends on LLM/agent output |
Tracks combine (e.g. a billing webhook in a multi-tenant SaaS that calls an LLM = core +tdd +saas +ai).
The engine: CLI and/or MCP (both local, zero-cost, no CI)
Do the mechanical steps with the bundled engine instead of hand-editing files. Two equivalent ways:
- CLI (works anywhere):
node cli/dev-spec.js <command>(ordev-spec <command>if on PATH). - MCP (if your tool speaks MCP): the
spec-drivenserver exposes the same operations as tools.
Key operations (CLI form):
dev-spec classify "<feature description>" # recommend tracks (multilingual, weighted)
dev-spec init [tracks...] [--lang en|pt|es] # scaffold .specs/steering (incl. constitution.md); --lang sets the project default
dev-spec steering <file> [--lang] # one steering file from its template (constitution.md, tech.md, …)
dev-spec create "<name>" [tracks...] [--lang] [--summary "…"] # scaffold the feature (no tracks → auto-classify; inherits project lang)
dev-spec status [feature] | list # progress, phase, tracks
dev-spec clarify <feature> # surface requirement gaps before design
dev-spec doctor <feature> # health-check → ready to advance? (exit 1 on FAIL — scriptable)
dev-spec ears <feature|file.md> # lint EARS (SHALL/DEVE/DEBE, IDs, vague words)
dev-spec trace <feature> # AC ↔ task ↔ test ↔ code (_Implements:_, phantom refs)
dev-spec next <feature> [--batch] # next task (--batch: + the [P] tasks that can run beside it)
dev-spec done <feature> <n> --run # run the task's _Verify:_ command and record the evidence (failure → stays open)
dev-spec bugfix "<name>" [--summary "…"] # bugfix flow: reproduce → root cause → regression test → fix
dev-spec finish <feature> [--write] # blockers + fresh checks + merge summary from the spec chain (merge locally; no PRs)
dev-spec brief <feature> [n] [--write] # self-contained brief for one task (ACs + tests resolved, DoD)
dev-spec approve <feature> <phase> # record an approval gate
dev-spec roadmap # multi-feature roadmap: %, dependencies, cycles
dev-spec depend <feature> [deps...] # declare dependencies / order (rejects cycles)
dev-spec scan [path] / dev-spec coverage # brownfield: inventory existing code + spec coverage
dev-spec evals <feature> [--dry-run] # run local eval harness (+ai; your API key)
dev-spec mcp-config [client] # print MCP config for your tool
The pipeline (Spec mode)
For anything beyond a quick fix (Vibe mode = just do it, no artifacts):
- Classify —
dev-spec classifyto seed tracks; confirm againstreferences/classification-matrix.md; write.specs/<feature>/classification.md. Get user approval of the track set. - Requirements —
dev-spec createscaffolds; write EARS criteria with stable AC IDs; rundev-spec earsto lint. Add track-specific ACs (tenant isolation for +saas; quality/safety/cost for +ai). Approve. - Design — base sections + the mandatory sections of the active tracks (5 for +saas, 10 for +ai). The scaffold marks each with a
> **TODO**sentinel; replace it with real content. No blank mandatory sections. Approve. - Test/Eval plan — +tdd: enumerate tests mapped to AC IDs. +ai: golden/adversarial/regression sets + thresholds + baseline. Approve.
- Failing tests / eval harness — +tdd: write tests, all red for the right reason (hard gate). +ai: deterministic tests + runnable eval harness + baseline. No implementation before this passes.
- Tasks — ordered, traceable; markers
_Requirements:_always,_Makes green:_(+tdd),_Emits metrics:_(+saas),_Affects evals:_(+ai). Rundev-spec trace— every AC must map to a task. - Execute — per task: implement-and-test (core) / red→green→refactor (+tdd) / prompt-iteration gated on eval delta (+ai).
dev-spec brief <feature>gives you the task with its ACs and tests already resolved — handy to focus, or to hand one task to another agent. Mark done withdev-spec done <feature> <n> --run— evidence before claims: the task's_Verify:_command runs and its result is recorded; a failure leaves the task open. Close the feature withdev-spec finish. (In Claude Code,/executeTask --subagentsruns an implementer + reviewer subagent per task — seereferences/subagent-execution.md; tools without subagents run inline.) Before "done": load test + observability (+saas), cost + safety validation (+ai).
At each phase boundary, run dev-spec doctor <feature>; only advance when it reports
readyToAdvance. Record sign-off with dev-spec approve <feature> <phase>.
Non-negotiables
- No implementation without approval at each gate.
- Traceability end-to-end: code → tasks → (tests/evals) → design → requirements → need.
- Mandatory track sections are mandatory — an honest "not needed because X" is fine; blank is not.
- Everything is local. No GitHub Actions, no paid CI, no pull requests — integrate by merging locally. Tests/load/evals run in the user's own env.
See skills/dev-spec-driven/references/ for EARS, scale, eval, safety, and prompt-engineering guides.
