Imported from Chargeuk/codeInfo2 (
AGENTS.md). Install upstream withnpx skills add Chargeuk/codeInfo2. Copyright stays with the author.
Agent Workflow Guide
Purpose
Use this file as the repository-specific operating guide for agents working in this repo. Keep the same behavior and standards described here even when the exact task changes.
Instruction Priority
- Follow the user’s direct request.
- Follow this repository guide unless the user explicitly asks to change repo workflow or policy.
- Prefer explicit repository facts and current code over assumptions.
- If required context is missing and can be gathered from the repo or tools, gather it before asking the user.
- If required context cannot be retrieved, ask the user only for the missing piece.
Session Start
Re-read this file at the start of each session. Assume it may have changed since the last context window.
Working with Planning Files
- Do NOT read whole planning files directly unless the user asks for specific plan content and the required information cannot be obtained from the Python helpers in
$CODEINFO_ROOT/scripts. Planning files can be very large, so query their structured summaries first and read only the smallest relevant section as a fallback. - Run these helpers from the target repository root so their default handoff paths resolve correctly. Prefer the default compact JSON output and add detail flags only when the task requires them.
- Use
python3 "$CODEINFO_ROOT/scripts/plan_status.py"for the common case. It resolves the active plan fromcodeInfoStatus/flow-state/current-plan.jsonand reports the selected task, task counts, completion state, unchecked subtasks or testing, and live blockers. - Narrow
plan_status.pyinstead of opening a plan:python3 "$CODEINFO_ROOT/scripts/plan_status.py" --task-number <number>inspects one task.python3 "$CODEINFO_ROOT/scripts/plan_status.py" --plan <path-to-plan.md>inspects a specific plan without changing the handoff.python3 "$CODEINFO_ROOT/scripts/plan_status.py" --include-tasksreturns every task summary; use it only when the compact default is insufficient.
- Use
python3 "$CODEINFO_ROOT/scripts/plan_sections.py"when an agent needs plan prose. It returns only requested sections with exact line ranges and completeness indicators:--profile implementation --task currentreturns the current task's implementation packet.--profile automated-proof --task current,--profile manual-proof --task current, and--profile blocker-repair --task currentreturn purpose-specific task packets.--profile story-scope,--profile review-scope,--profile review-tasking,--profile testing-audit, and--profile closeoutreturn bounded cross-task or story context.--profile review-evidencereturns the story contract, repository scope, compact task index, and final-task proof packet;--profile review-findingsreturns story scope plus the task index for targeted review expansion.- Review profiles expose available story headings and task-section names without their prose so agents can request custom sections deliberately.
--task-number <number> --section <heading>and--story-section <heading>request one additional named section without opening the complete plan.
- Use
python3 "$CODEINFO_ROOT/scripts/story_workflow_status.py"for the combined story view, including plan completion, repository scope, and review-loop state. Add--include-tasksonly when per-task detail is necessary. - Use the focused read-only helpers when only one answer is needed:
python3 "$CODEINFO_ROOT/scripts/check_current_task_handoff.py"checks whether the persisted current-task handoff is still valid.python3 "$CODEINFO_ROOT/scripts/plan_blocker_status.py"reports blocker status for the selected task; it also accepts--task-number <number>or--plan <path-to-plan.md>.python3 "$CODEINFO_ROOT/scripts/questions_section_status.py"reports whether the active plan's Questions section requires attention.python3 "$CODEINFO_ROOT/scripts/manual_testing_guidance_status.py"extracts story-level manual-testing guidance; pass--task-number <number>for one task.
python3 "$CODEINFO_ROOT/scripts/select_current_task.py"writes the current-task flow-state handoff. Use it only when the workflow requires selecting or refreshing that handoff, not for a read-only status query.- If a helper's interface is unclear, run it with
--helpbefore reading its source or the planning file. - When the Python helpers do not expose the required detail, use shell tools to locate and print only a bounded section. Use this fallback order:
wc -l <plan.md>checks the file size before selecting a reading strategy.rg -n --max-count <count> '<heading-or-term>' <plan.md>finds relevant line numbers without printing the file. Add-C <lines>,-A <lines>, or-B <lines>only for small, bounded context around a match.sed -n '<start>,<end>p' <plan.md>prints an exact line range afterrg -nidentifies its boundaries.awk '/<start-heading>/{show=1} show{print} /<end-heading>/{exit}' <plan.md>extracts a heading-delimited section when line numbers are inconvenient. Choose a specific end-heading so output cannot run to the end of the file accidentally.head -n <count> <plan.md>ortail -n <count> <plan.md>reads a bounded introduction or ending only when the needed content is known to be there.git diff --unified=<lines> -- <plan.md>inspects only uncommitted plan changes with limited surrounding context.
- Use
jqto narrow JSON emitted by the Python helpers before requesting more plan text, for examplepython3 "$CODEINFO_ROOT/scripts/plan_status.py" | jq '{selected_task, story_complete, tasks_with_live_blockers}'. - Do not use
cat, an unboundedsed/awkrange, or a broad recursive search to read a large planning file. Start with the smallest likely query and expand the line range or context only when the first bounded result is insufficient.
Documentation Sources
- When working in React, use the MUI MCP tool for all Material UI references.
- For any other API or SDK, consult current documentation via the Context7 MCP tool.
Harness Path Contract
CODEINFO_ROOTalways means the harness repository root for this workflow, not the target repository root and not the currentworking_folder.- Use
"$CODEINFO_ROOT/..."only for harness-owned assets such asscripts/**,codeinfo_markdown/**,flows/**, and similar workflow support files. - When you need files from the target repository being worked on, use the selected repository path,
working_folder, repo-relative paths, or explicit absolute paths instead ofCODEINFO_ROOT.
Research And Documentation Order
- For questions about this local repository, use the
code_infoMCP tool first. - After
code_info, inspect the local code directly with repository search and file reads as needed. - For Material UI questions, use the MUI MCP tool first.
- For non-MUI libraries, SDKs, or API documentation, use the Context7 MCP tool first.
- For external GitHub repository documentation or architecture questions, use the DeepWiki MCP tool when it is relevant to the task.
- Do not force a single global tool order across all tasks. Choose the first tool based on whether the task is about the local repo, MUI, another library or SDK, or an external GitHub repository.
Flow Design And Agent Handoffs
- Follow KISS. Prefer the simplest reliable flow and avoid unnecessary handoffs, schemas, validators, intermediate files, and duplicated state.
- Prefer continuing from persisted flow state, immutable flow input, or the current repository state over transferring information between agents through files.
- When a file or artifact is necessary, the producing agent owns making it as compliant, complete, and self-describing as possible. State its purpose, provenance, outcome, remaining uncertainty, and whether the result is complete, partial, or unavailable so a reader can understand it without prior knowledge of an exact format.
- Consumers must interpret files and artifacts semantically and make a best effort. Missing, malformed, incomplete, contradictory, or unexpectedly formatted data is an evidence limitation; salvage every trustworthy fact and do not reject the input merely because its format is imperfect.
- A missing or imperfect applicable or expected handoff must trigger best-effort recovery and visible accounting. If trustworthy evidence cannot establish a complete result, produce an honest partial or unavailable result; the limitation must not by itself fail the agent turn or stop the parent flow. A later artifact that trustworthy earlier evidence has established as deliberately inapplicable requires no warning or placeholder; preserve its skip reason in the next applicable audit, disposition, or outcome. Continue with whatever safe, useful work remains.
- When configuration defines independent list or batch entries, validate entries independently and continue with every valid, unambiguous entry whenever partial execution is safe. Apply a documented deterministic policy to duplicates or recoverable conflicts, and log a visible secret-free warning for every malformed, conflicting, discarded, or unavailable entry. Never guess a safety-critical identity, endpoint qualification, credential, permission, or missing required value; isolate the unsafe entry, and fail the whole operation only when no safe scope remains or atomicity is explicitly required.
- Best effort does not permit invented evidence, unsafe path guesses, destructive action without authority, or a false claim of success. When a safety-critical identity, path, permission, or input cannot be established, take a safe no-op for the affected action, report the limitation clearly, and allow the surrounding flow to continue where possible.
- Apply the detailed recovery and artifact-consumption rules in
codeinfo_markdown/shared/review-artifact-handoff.mdto review artifacts.
Output Contract
- Keep responses concise and task-focused.
- Include file references when explaining code or repo instructions.
- State clearly when tests, builds, or validation steps were not run.
- Do not invent repository state, tool output, or test results.
- Before substantial work, send a short progress update describing the next action.
- After substantial work, report what changed and how it was validated.
Working Through Story Plans
When working from a file in ./planning that is NOT a pr-summary file, update the plan continuously as implementation progresses.
- Mark each subtask complete immediately when that subtask is implemented.
- Mark each testing step complete immediately when that testing step is performed.
- Change each completed checkbox from
[ ]to[x]at the point of completion. - Do not batch multiple checkbox updates later.
- After marking a subtask or testing step complete, add a brief point to that task’s
Implementation Notessection. - Each implementation note must briefly state:
- what was done;
- any issue that had to be overcome;
- or why the step is blocked.
- This
Implementation Notesupdate is plan-maintenance, not a separate subtask or testing item, and should not be pre-modeled as a future-dependent checklist entry. - Every unchecked testing checkbox is mandatory blocking work, even if its wording says
optional,if needed,if targeted diagnosis is needed, or similar. - Do not model optional or failure-only diagnostic commands as unchecked testing checkboxes. Keep them in prose or inside the failure-diagnosis text of a mandatory testing step instead.
Branching And Phase Flow
- Create or reuse a feature branch for each story using
feature/<number>-<short-description>. - Create that branch from the currently checked out location unless the user instructs otherwise.
- Work only within that branch until the story is complete and working.
- Each commit message must start with
DEV-[Number] -. - Each commit message body should contain 4 or 5 sentences explaining what changed and why.
Build Workflow
Wrapper-First Rule
- Use build wrappers as the default build path.
- Use raw underlying commands only when wrapper maintenance or diagnosis requires them.
- Wrapper logs are the source of full diagnostic detail.
Build Wrappers
npm run build:summary:serverbuilds the server workspace with compact summary output. Full log:logs/test-summaries/build-server-latest.log.npm run build:summary:clientruns the client workspace typecheck as a pre-build gate and then the client build with compact summary output. Full log:logs/test-summaries/build-client-latest.log.npm run typecheck:summary:clientruns the client workspace typecheck with compact summary output when direct TypeScript diagnosis is needed without the build phase. Full log:logs/test-summaries/typecheck-client-latest.log.npm run compose:build:summaryruns Docker Compose build with compact summary output. Full log:logs/test-summaries/compose-build-latest.log.
Summary Wrapper Output Contract
- Summary wrappers emit heartbeat/final guidance fields on wrapper stdout:
timestamp,phase,status,log_size_bytes,agent_action,do_not_read_log, and a finallogpath. - If
agent_action: wait, the wrapper is still healthy and running. - While the wrapper continues to emit heartbeat updates at least about every 2 minutes, keeps
do_not_read_log: true, and shows ongoing progress such as growinglog_size_bytes, you MUST keep waiting regardless of total elapsed wall-clock time. - Do not treat elapsed time by itself as evidence that the wrapper is hung.
- Do not read the saved log while
do_not_read_log: true. - If
agent_action: skip_log, the wrapper finished with a clean success. Do not read the saved log unless the user asks or later work specifically needs it. - If
agent_action: inspect_log, the wrapper ended with warnings, failure, or ambiguous parsing. Open the saved log and diagnose from there.
Build Failure Diagnosis
- Run the relevant wrapper.
- If the wrapper is still emitting timely
agent_action: waitheartbeats and ongoing progress, continue waiting and do not interrupt it solely because it has been running for a long time. - Capture the log path from the wrapper summary output if the wrapper ends with
agent_action: inspect_logor otherwise fails unexpectedly. - Open the log file and inspect the failing command output.
- Fix the failing dependency, config, or code.
- Re-run the same wrapper.
- If a Docker or Compose build fails due to a likely transient network or cache issue, retry once before deeper investigation.
Run Workflow
Wrapper-First Rule
- Compose wrappers are the default way to start and stop the AI-agent testing or automation stack.
- For manual story proof and close-out testing, prefer the checked-in main
codeinfostack fromdocker-compose.yml, notcodeinfo:local. - Do not run any
*:local:*commands from this agent. - These wrappers centralize env-file handling and Docker socket or runtime compatibility through
scripts/docker-compose-with-env.sh.
Run Wrappers
npm run composebuilds and starts the testing stack.npm run compose:buildbuilds the testing stack images.npm run compose:upstarts the testing stack when images are already built.npm run compose:downstops the testing stack.npm run compose:logstails testing stack logs.
Preferred Start Sequence
- Run
npm run compose:build. - Run
npm run compose:up.
Manual-proof notes:
- Treat the main stack at
http://localhost:5001andhttp://localhost:5010as the supported human/manual-testing surface unless the active plan says otherwise. - The checked-in main stack mounts its editable proof agent catalog from
manual_testing/codeinfo_agentsandmanual_testing/codex_agents; when later manual proof needs dedicated warning, fallback, duplicate-root, provider-specific, or limited-capability cases, prefer adjusting thosemanual_testingroots rather than disturbing thecodeinfo:localdevelopment catalog. - Manual testing may be skipped only when repository-owned guidance in
AGENTS.mdorcodeinfo_markdown/repository_information.mddefines a skip condition and that condition is actually being hit during proof. For this repository, missing provider login or auth state that requires human-controlled two-factor authentication is an allowed skip for the affected auth-dependent surface. Autonomous manual proof must not attemptRe-authenticatein that case, must not reopen or fail the task by itself, and must not generate implementation work by itself.
Shortcut:
npm run composeis the equivalent build-plus-up sequence.
Stop Sequence
- Run
npm run compose:down.
Run Failure Diagnosis
- Re-run the relevant compose wrapper and capture the terminal output.
- Tail logs with
npm run compose:logs. - Fix the failing container, config, or env issue.
- Re-run the same wrapper.
Repository-Owned Test Stack Reclamation
- Manual and automated testing agents may stop a pre-existing or stale Docker or Compose stack when current repository evidence proves that it belongs to a repository they are permitted to test and is the documented stack required by the current testing step.
- Establish ownership from repository-supported wrappers, Compose configuration, and Compose metadata or labels. An occupied port or container name alone is not sufficient.
- Reclaim the stack with its repository-supported shutdown wrapper, then continue the documented startup and proof lifecycle. The stack does not need to have been started by the same agent or flow step.
- Do not interrupt a healthy stack in the middle of the current startup, test, and shutdown lifecycle.
- If repository ownership or testing applicability remains uncertain, do not stop the stack; report the conflict honestly.
- These permissions apply to repository test stacks only. The protected local development stack rules below still take precedence.
Local Stack Safety
- If
docker-compose.local.ymlservices are running, assume they may be hosting the current Codex or manual-testing session. codeinfo:localis the live development stack, not the default close-out proof stack. Keep its bind-mountedcodeinfo_agentsandcodex_agentstrees untouched unless the user explicitly wants development-stack changes there.- Do not run
npm run compose:local:down,docker compose -f docker-compose.local.yml down, or otherwise stop or removecodeinfo2-*-localcontainers unless the user explicitly instructs you to do so. - Do not restart or clean up the local stack just because it appears stale. If a task seems to require switching away from the local stack, stop and ask first.
- Reason: the local compose stack may be the live runtime backing the current agent session, browser tooling, or proof flow; taking it down can kill the session and interrupt work in progress.
Test Workflow
Wrapper-First Rule
- Use test wrappers as the default test path.
- Use raw underlying commands only when wrapper maintenance or diagnosis requires them.
- Wrapper logs are the source of full diagnostic detail.
- When a task requires running the full automated test suite across client, server, and e2e surfaces, use
npm run test:summary:all:parallelas the required all-tests wrapper.
Parallel-Safe Test Authoring
- Observe asynchronous results before triggering the operation that can produce them. Create event, response, abort, or completion waiters before sending a request, publishing a message, clicking an action, or starting the background work.
- For server WebSocket tests, use
connectWs,subscribeConversationAndWaitReady,waitForEvent,waitForClose, andcloseWsfromserver/src/test/support/wsClient.ts. Do not attach a raw WebSocket listener after sending a message unless an earlier listener is already buffering every relevant event. - Prefer deterministic readiness boundaries such as deferred promises, explicit gates, emitted events, or observable state. Do not use
sleep,delay, orsetTimeoutmerely to give work time to finish before an assertion. - When polling is unavoidable, use
waitForTestConditionfromserver/src/test/support/testTimeouts.ts, or an equivalent surface-specific bounded helper, with a useful failure description and a configured timeout. Fixed delays are permitted only for deliberate mock pacing, retry simulation, or timeout behavior, never as the correctness boundary. - Prove that something has not happened by holding execution at a known deterministic gate and inspecting state there. Do not sleep for an arbitrary interval and then make a negative assertion.
- Cancellation-aware test doubles must handle an already-aborted signal before waiting, register abort listeners with
{ once: true }, and clean up non-one-shot listeners. If an asynchronous boundary can occur between checking and registration, use an idempotent completion handler and recheck the signal after registration. - Do not introduce unscoped shared mutable state. Use the scoped test environment and override helpers instead of direct
process.envmutation, and restore mocked globals such asconsole,Date.now,fetch, provider singletons, and dependency overrides infinallyor cleanup hooks. - Isolate external resources. Prefer port
0for test servers,fs.mkdtempor the provider-home harness for filesystem state, and unique test-owned identities for Compose projects, databases, collections, and similar shared resources. - Await all asynchronous cleanup. Sockets, servers, timers, child processes, temporary directories, subscriptions, client pools, registry entries, and intentionally detached promises must be owned by the test and settled in
finally,afterEach, orafterAll; do not rely on process exit. - Route timeouts through
resolveConfiguredTestTimeoutMs,resolveClientTestTimeoutMs, orresolveConfiguredE2eTimeoutMsas appropriate. Do not introduce short hard-coded timeout assumptions that become invalid under CPU saturation. - Run the smallest applicable summary wrapper first. New or changed concurrency-sensitive tests must then be validated with
npm run test:summary:all:stress; passing repeatedly in isolation is not sufficient proof of parallel safety.
Safe event-ordering pattern:
const resultPromise = waitForEvent({
ws,
predicate,
describe: () => 'expected result event',
});
sendJson(ws, request);
const result = await resultPromise;
Prohibited timing-dependent pattern:
sendJson(ws, request);
await delay(100);
expect(result).toBeDefined();
Test Wrappers
npm run test:summary:clientruns the client test suite with compact summary output. Full log:test-results/client-tests-<timestamp>.log. JSON:test-results/client-tests-<timestamp>.json.npm run test:summary:server:unitruns the servernode:testunit and integration suites with compact summary output. Full log:test-results/server-unit-tests-<timestamp>.log.npm run test:summary:server:cucumberruns the server cucumber feature suites with compact summary output. Full log:test-results/server-cucumber-tests-<timestamp>.log.npm run test:summary:e2eruns the e2e flow with setup, build, tests, and teardown. Full log:logs/test-summaries/e2e-tests-latest.log.npm run test:summary:server:parallelbuilds the server workspace once, then runs the server unit and cucumber wrappers in parallel with--skip-build.npm run test:summary:all:parallelbuilds reusable client, server, and e2e compose artifacts first, then runs the client, server unit, server cucumber, and e2e wrappers in parallel with shared-build skip flags.npm run test:summary:all:stressuses the same allocation as the parallel wrapper, then assigns otherwise-unused cores to the server unit suite to expose timing and isolation defects under higher concurrency.npm run test:summary:client:parallelis a convenience validation path that runsbuild:summary:clientand thentest:summary:client; it belongs to the parallel workflow family even though it is not a multi-harness fan-out by itself.- These wrappers do not have a fixed failure time budget.
- As long as a wrapper continues to emit healthy
agent_action: waitheartbeats at least about every 2 minutes and shows ongoing progress such as growinglog_size_bytes, you must keep waiting no matter how long the run takes. - Standalone wrappers remain the self-contained default for diagnosis. Use the new
*:parallelcommands for batch validation when you want shared prebuilds and cross-harness parallelism. - Treat
npm run test:summary:all:parallelas the canonical batch-validation wrapper whenever repo instructions, a plan, or a task says to run "all tests", "the full automated suite", or equivalent full-suite wording.
Targeted Test Runs
- Client Jest supports
--file,--subset, and--test-name. - Server
node:testwrapper supports--file,--test-name, and--skip-build. - Server cucumber wrapper supports
--tags,--feature,--scenario, and--skip-build. - E2E Playwright wrapper supports
--file,--grep, and--skip-compose-build. - For final validation, run the full relevant summary wrapper without targeted args.
- When final validation must cover the entire automated repo test surface, that full wrapper is
npm run test:summary:all:parallel.
Test Failure Diagnosis
- Run the relevant wrapper.
- If the wrapper is still emitting timely
agent_action: waitheartbeats and ongoing progress, continue waiting and do not interrupt it solely because it has been running for a long time. - Capture the log path from the wrapper summary output if the wrapper ends with
agent_action: inspect_logor otherwise fails unexpectedly. - Open the full log file and locate the failing test block.
- Fix the failing code, config, or dependency.
- Re-run targeted wrappers for diagnosis as needed.
- After fixes, re-run the full relevant summary wrapper without targeted args.
