Imported from xdanny/multiturn-sql-finetuning (
docs/AGENTS.md). Install upstream withnpx skills add xdanny/multiturn-sql-finetuning --skill docs. Copyright stays with the author.
Docs AGENTS
This subtree owns the research narrative, roadmap, and small checked-in documentation artifacts for the SQL finetuning program.
Primary responsibilities:
- Keep blog-derived ideas grounded in runnable repo steps: direct SQL controls, structured query briefs, semantic context, metric DSL, behavior recovery, and hosted or BIRD-style comparisons.
- Separate smoke-run viability from scored evaluation and benchmark claims.
- Keep claim language human-readable, specific, and tied to visible artifacts.
- Make docs point to the exact code, data artifact, command, manifest, or comparison file they rely on.
Rules:
- Do not describe a method as better until the required comparison manifest exists and shows the right same-row delta.
- Do not turn temporary outputs into checked-in evidence. Runtime outputs belong
under
outputs/orresults/; small canonical inputs belong underdocs/data_artifacts/. - Do not expose future turns, reference SQL, gold plans, gold metric DSL, expected rows, or repair labels as production-style model inputs.
- When adding a method doc, state the control arm, primary metric, supporting metrics, leakage boundary, and comparison artifact.
- Prefer short prose and compact bullet lists over large markdown tables.
- If a doc cites CoSQL, SParC, BIRD, or BIRD-Interact, explain the role of that benchmark in this repo instead of assuming the reader already knows it.
- Use
uv run ...in documented Python commands.
When editing here, inspect:
docs/research_roadmap.mddocs/current_research_inventory.mddocs/data_artifacts/README.mdREADME.md
