Imported from ziweiwu/agent-commander (
AGENTS.md). Install upstream withnpx skills add ziweiwu/agent-commander. Copyright stays with the author.
Working on agent-commander
Instructions for Claude Code and any other agent working in this repository. Read this before changing anything.
What this is
A local web dashboard over every Claude Code session on the machine — status, folder, what each one is doing — with a terminal you can answer a blocked agent from, including from a phone over Tailscale.
It is an observer of somebody else's work in progress. That is the whole design pressure: the sessions on screen are real, and this app is never allowed to disturb one.
The fleet has one rendering: FleetList, the grouped card list. There were
two — a ForestView drew each session as a family on a shared time axis — and
the split cost more than it returned. Every property that had to be true of the
fleet had to be proved twice, and the forest's own question ("is anything in
this family still moving?") turned out to be answerable on the card itself, from
the same delegation trees the forest was reading. So the trees moved onto the
card: AgentCard rolls each agent's delegates into a line and opens them in
DelegationTree, and the reasoning that has to survive that move is INV-13 and
INV-15. The pure parts — what a card may claim about delegates, the two lengths
of its activity trail, and what its status rail may assert — live in
src/web/lib/delegation.ts, src/web/lib/trail.ts and src/web/lib/status.ts
so they can be tested without a DOM.
The card's face is a status rail plus two lines: a fixed glyph gutter
answering "which of these needs me" by position rather than by reading, then
the name and the age, then a context line. status.ts holds the two channels
the rail is built on and keeps them apart — what a session is doing, and
whether this app can still reach it, which is where attachBlockedReason
lives now. INV-11 carries the rules; the short version is that the hand is
never drawn over an inferred status and the working arc turns without ever
filling.
Two languages, and which is which
The server is Rust, in rust/. The browser app is TypeScript, in src/web/.
rust/src/types.rs is the wire contract, and src/shared/wire.ts is
generated from it by npm run gen:types — types and the option lists both,
because ALLOWED_KEYS, MODEL_ALIASES and friends are what make "the server
validates against them and the browser offers exactly them" true, and a
type-only export would have left that half to drift. Edit the Rust, run the
script, commit both. types::tests::the_checked_in_wire_contract_is_current
fails npm test on a checkout where they disagree, so forgetting the second
step is a red gate rather than a browser bug. src/shared/types.ts is the
hand-written remainder: it re-exports wire.ts and adds the two union types
the browser derives from the lists. The derive is cfg_attr(test) and ts-rs
is a dev-dependency, so none of it reaches the release binary.
There is no src/server/ any more. It was ~6,600 lines of TypeScript and it is
preserved, working, on the old-node-backend-branch branch. Reach for it
when you want to know what the old code did; do not reach for it to copy code
back.
The port is verified by three things rather than by reading:
rust/tests/golden/*.json— real responses captured from the Node server on the mock fleet.mock.rs's tests assert the Rust bytes still match them, so a drift in any wire field fails a unit test rather than a browser.- The 233 Playwright tests, unchanged. They drive HTTP and WebSocket and never imported the server, so they arbitrate the port without knowing it happened.
cargo test, which carries the invariant numbers the same way the TypeScript did:cargo test --manifest-path rust/Cargo.toml inv13.
What moved and what did not: every module in src/server/*.ts has a
same-named rust/src/*.rs, except that tmux-client.ts gained
tmux_source.rs/tmux_agents.rs beside it, agent-kinds.ts became
agent_kinds.rs on the server side while staying TypeScript for the browser,
and cli.ts split into main.rs (entry, statusline install) and options.rs
(argument parsing, the port guards, INV-3's bind refusal).
The two constraints everything else follows from
ARCHITECTURE.md opens with these, and says explicitly that everything else is
downstream of them.
INV-1 — no tmux client this app creates may affect the size of a pane. It is
why the Attach view is capture-pane polled, diffed and replayed into xterm.js
rather than a pty. A pty means a client, a client has a size, and under
window-size latest with aggressive-resize on a browser-shaped client reflows
the pane a working agent is drawing into. The browser's own width only ever sets
a CSS transform.
INV-2 — nothing reaches a live agent except from an explicit user action. No retries, no auto-send, no replay on reconnect. It is why the client carries four separate duplicate-suppression mechanisms rather than one, why a paste is staged through a file instead of a command line, and why every control action is verified by reading the transcript back rather than assumed to have worked.
A change that makes either of those less true is the wrong change however
convenient it looks. ARCHITECTURE.md §"The two constraints everything else
follows from" carries the reasoning; INVARIANTS.md carries the contract.
Before you say it works
Run these. If you did not run them, say so explicitly rather than implying success.
npm run typecheck
npm run lint
npm test # 1504 tests: 691 Rust (the server) + 813 vitest (the web app)
npm run build # vite bundle, then `cargo build --release`
npm run e2e # 406 end-to-end tests, five projects: desktop/tablet/phone on
# Chromium, and phone/tablet again on WebKit. Two mock
# servers: the fixture fleet on 4599 and `--mock-empty`
# on 4598, which `e2e/empty.spec.ts` alone points at.
# E2E_PORT and E2E_EMPTY_PORT move them if those are taken
npm run audit # contrast, a11y, task flows, device layouts — needs a server
npm run audit:workspace # the work-surface bar: >=80% of the viewport is transcript
# plus composer at desktop/laptop/tablet, measured in a real
# browser against --mock. Not in `npm run audit`, because it
# asserts a product target rather than a correctness property.
# PORT/BASE move the server (4400 by default, and it refuses
# 4317 outright), AGENT picks the fixture, BAR moves the
# threshold — the same BASE/PORT the other audit scripts read
npm run qa # randomised exploration, deterministic per seed. It does
# NOT start its own server: put a `--mock` one on 4500
# first, or all twelve seeds fail with a connection
# refused that reads as twelve findings
npm run verify:inv1 # attaching never resizes a real pane — server must be running
ARCHITECTURE.md §"How it is checked" is the table to read before touching any
of them: five gates, each answering a question the others cannot, and the note
on why the last three stay local habits rather than CI gates.
npm test and npm run e2e run on every push and pull request
(.github/workflows/ci.yml), which installs a stable Rust toolchain (cached
with Swatinem/rust-cache) alongside both Chromium and WebKit. The
second engine is not redundancy: every browser on iOS is WebKit, and for this
app's first five releases every "iPhone" and "iPad" result it produced was
Chromium wearing an iPhone user-agent. ENGINE=webkit npm run audit:mobile
runs the device audit on WebKit too. The rest are local.
Everything there builds on stable, so the floor job also checks the server
on the toolchain rust-version in rust/Cargo.toml names. That number is a
promise to anyone building from source, and the lockfile once moved past it
with every gate green. A cargo update that needs a newer compiler fails
there: run it with CARGO_RESOLVER_INCOMPATIBLE_RUST_VERSIONS=fallback, or
raise rust-version and the handbook's line together.
The first three also run as a Stop hook, so a turn that leaves the tree
failing is refused rather than summarised. The hook lives in the harness
plugin and reads its list from .claude/gates.json here — watch
pathspecs, gates as [name, argv] pairs, and a per-gate timeout. It skips
when nothing under src/, test/, scripts/ or the manifests has changed, and
blocks at most once per tree state so a failure it cannot fix never traps the
session.
build, e2e, audit:*, qa, verify:inv1 and app are deliberately
not in that list. Measured on this machine: typecheck 5.0s, lint 0.9s, test ~40s,
build 3.4s — the three that are in it already cost ~46s, and a Stop hook slow
enough to resent is one that gets deleted. The others need a browser, a running
server, or a live tmux session with a real agent in it, none of which a hook can
assume. CI and a deliberate local run own them. app is the same argument in a
different key: it is macOS-only, it needs a full build first, and what it guards
is a local convenience rather than anything the app does — its three cheap
checks already ride in npm test, and scripts/mac-app/ is covered by the
existing scripts/ watch entry.
watch entries are git pathspecs, not prefixes — git matches whole path
components, so the manifests are spelled out individually. scripts/ is on the
list because three generated files are held to their generators by tests there,
and because scripts/cargo.sh is how every gate reaches the Rust toolchain.
rust/ is on it for the obvious reason: it is the server.
Three things are generated, and each has a test that says so
The pattern is the same every time: the generator carries the drawing and the reasoning, the output is committed, and a test re-runs the generator and fails a checkout where the two have drifted. Edit the generator, never the output.
| Generator | Output | Held by |
|---|---|---|
scripts/gen-themes.py |
src/web/styles/tokens.css |
test/scheme.test.ts |
scripts/gen-ui-icons.py |
src/web/lib/icon-paths.ts |
test/icons.test.ts |
scripts/gen-icons.py |
src/web/public/assets/icon-*.png |
test/mac-app.test.ts |
The last two are different scripts and the names are one word apart.
gen-icons.py draws the application icon — the PWA PNGs and the macOS
.iconset, one picture of three lanes. gen-ui-icons.py draws the
small control faces inside the app. They share a prefix and nothing else, and
the collision has already cost one accidental overwrite; each file's docstring
opens by saying which one it is.
Anything a gate shells out to must go through scripts/cargo.sh, never bare
cargo. ~/.cargo/bin is put on PATH by a line in your shell profile, so it
is present in a terminal and absent in every non-interactive shell — npm
scripts, git hooks, and the Stop hook that runs these gates. Bare cargo in a
package script passes locally and fails with cargo: command not found for
everyone and everything else, which is exactly how it was first written here.
Ports, and why there are three
| Port | What runs there |
|---|---|
| 4317 | Production. Real agents. npm start, npm run serve, the installed binary. |
| 4400 | Development. npm run dev, npm run mock, and what the audit scripts target. |
| 4500 | qa-sweep.sh. |
Never point a fuzzer, an audit or a review agent at 4317. It drives real
agents, and anything that types into whatever it finds will type into someone's
session. qa-sweep.sh refuses that port outright and --mock on it is rejected.
npm run mock serves a deliberately awkward fixture fleet — fifteen sessions,
five sharing a home directory, one name too long for its card, five never
prompted, all three shapes an agent blocks
on (a question with options, a plan awaiting approval, a tool awaiting
permission — INV-16's three) plus a two-question set whose pane moves under
the keys the answer card sends (mock-set, reserved for
e2e/question-set.spec.ts the way mock-idle-db is reserved for /clear),
one whose pane has exited, so the Attach tab's
dead-pane notice is a thing you can look at, and one plain terminal — which is
not on screen at rest, because terminals are out of the fleet's scope until
the chip admits them, and a filter with nothing behind it cannot be looked at. --mock-empty serves the
same server with no agents in it, which is the only way to see the
confirmed-empty screen rather than the loading one. The delegation trees behind them are
awkward on purpose too: a depth-3 chain, a delegate the user stopped, one node
in each of INV-13's three states, an orphan whose parent is not on disk, an
agent that has delegated nothing, and a CLI that cannot say either way.
Because the mock fleet runs the same server, routes and validation as the real
one (rust/src/sources.rs is the seam), a failure seen in mock mode is the
failure you would get for real.
Running it while you work on it
Three servers, and only one of them is yours to restart casually.
4400 is the one to develop against. npm run mock builds the bundle, builds
the server and serves the fixture fleet. Every audit script targets it, and a
failure seen there is the failure you would get for real, because mock mode
swaps only sources.rs. Restarting it is kill on whatever holds the port,
then run it again — there is nothing watching it.
4317 is production, and it is a launchd job. ~/Library/LaunchAgents/ com.ziweiwu.agent-commander.plist owns it with KeepAlive, so a plain kill
brings the old binary straight back and looks like your change did nothing.
Restart it properly:
npm run build # bundle, then the binary
launchctl kickstart -k gui/$(id -u)/com.ziweiwu.agent-commander
tail -2 ~/Library/Logs/agent-commander.log # it prints the URL
The token survives that: --token auto reads ~/.claude/agent-commander/token
rather than minting a new one (token_file.rs), so a link saved on a phone
keeps working across restarts, and tailscale serve needs no attention at all
— it proxies the port, and the port does not move. Check it with
tailscale serve status if something looks unreachable, but suspect the server
before the proxy.
A rebuild without that restart is the trap. npm run build:web rewrites
dist/web under a server that has been up for days, so the page is new and the
binary answering it is not. That pairing is ordinary here, which is why the
client learns whether the server beats rather than assuming it (INV-4) — but
anything needing a new route or a new wire field will simply not work until the
binary is replaced, with nothing saying why.
One browser, shared. The Chrome the DevTools MCP drives is a single
instance. A subagent that opens a page takes the selected tab with it, so a
screenshot taken while one is running can silently be of its page rather than
yours — the file on disk is fine and the picture is of something else. Check
document.querySelector('h1') before believing a render, or serialise: do not
drive the browser while a browser-driving agent is out.
The invariant contract
INVARIANTS.md numbers every property this app is built against, INV-1 through
INV-18, and each is greppable from a test name:
cargo test --manifest-path rust/Cargo.toml inv3 # the server's half
npm run test:web -- -t INV-3 # the browser's half
When you add behaviour worth relying on, add a numbered invariant and a test carrying its number. When you change behaviour, update the invariant in the same commit.
Things that have already bitten
-
Kiro CLI support was removed, deliberately. The app once listed Kiro sessions found through tmux — by session name or process name — as a degraded card with no conversation. Every flag on that row was a claim about another program's interface checked against no version, and it went stale for a whole major version before anybody looked. Rather than carry a second CLI it could only half-read, the fleet is Claude Code plus this app's own terminals;
agent_kinds.rshas two rows andtmux_agentsrecognises a pane only by the marker this app wrote on it. Do not re-add a kind by name or process matching without a reader for what it writes. -
The bridge writes two things now.
scripts/statusline-bridge.mjsspillsrate_limitstorate-limits.jsonas before and, per session, the context percentage and cost tosessions/<session_id>.json;usage.rsreads the second inside the enrichment pass.--install-statuslineis unchanged. -
Push goes out through
--notify, and only out.push.rsposts to an ntfy topic or a Telegram bot when an agent becomes waiting while no browser reports itself visible (INV-14, server-side);--notify-linkis this app's address on the phone so the push opens the card. It is off unless the flag is given, and there is no channel in mock mode to point a review agent at. -
binpoints at a launcher, not at the server. The server is a Rust binary, andbinhas to name something node can run, soscripts/launch.mjsfindsrust/target/release/agent-commanderand execs it. Source edits do nothing untilnpm run build:server. The launcher fails loudly when the binary is missing, and that is deliberate: every global install from 0.1.0 shipped a CLI that produced no output, opened no port and reported no error, and a silent launcher would be that bug again wearing a different hat. The published package now carries prebuilt binaries so the launcher has something to find — see Shipping it to npm below. -
The do-nothing install has shipped twice, and the fat package is the third attempt to stop it. First the
realpathSyncbug, where a symlinkedbinmade the CLI decide it had been imported and never callmain(). Then the port, wherefilesshippedrust/srcand no binary. Neither failed with a message; both installed cleanly, ran, and produced nothing. That history is whylaunch.mjsprints every path it tried and the host it is on before exiting 1, why the release publishes one package instead of a matrix that has a window where a platform package is missing, and whybuild-mac-app.pysmoke-tests the staged bundle. When you change anything on the path from a tag to an installed binary, the question to ask is not "does it work" but "how would I find out if it did not". -
A token in the URL cannot reach a subresource.
--token401'd the app's own bundle, becauseindex.html's<script>and<link>carry neither the token nor a header.GET/HEADunder/assets/skip the token gate and nothing else does — so nothing under that prefix may ever serve agent state. The same assumption broke the address bar: the router replaces the whole location, so navigation has to re-attach the token. -
The two gates are independent, and the origin one is never skipped.
permittedissame_origin_request(headers, &self.origin_names). A token no longer exempts a request from rebinding protection — it used to, which made the token the credential and the exemption at once.origin_namesis empty without a token, so a tokenless server answers to loopback alone; the Tailscale name buys nothing on its own, becausetailscale servehands every tailnet peer the same name. The token travels in the query string once and is traded for a cookie signed with it;announcemasks it unless--print-url. The origin gate compares the port as well as the name, and a real server keeps a token by default (--no-tokento opt out on loopback). SeeARCHITECTURE.md§"Where it is fragile" 6 and 6a before touching either gate. -
Development used to default to 4317. A fixture fleet on the production port is indistinguishable from your real one having vanished, and the composer on that page types into nothing.
-
Registry.changed()does not watch enrichment fields (registry.ts:320).activity,goalandmodelreach the browser only because the enricher callsnotify()itself. Forget that and the UI lags indefinitely with nothing raising an error. -
A running server outlives the bundle it serves, and an unknown route is not an error.
npm run buildrewritesdist/webunder a server that has been up for days, so a page newer than the binary answering it is what a rebuild ordinarily produces. A route that binary has never heard of is answered with the SPA shell —200 text/html, because the app is served from every path that is not an endpoint — sores.json()throws and the parser's words reach the user as the server's reason: "Could not open the terminal: The string did not match the expected pattern." Nothing 404s, nothing logs, and the feature reads as broken.spawnRequestchecks the content type and says what it actually means. Every new endpoint inherits that shape; the first thing to check when a new route "does not work" ispson the server that is answering it. -
A placeholder must never shadow the session it stands in for.
pendingannounces apending:card from the Claude registry, which is provider zero inCompositeSource; a terminal is found by the tmux sweep, which is provider one. On plain provider order the placeholder won the tmux session and the real terminal was dropped for the five minutes a placeholder may live — andpending's own retirement rule could not end it either, because that fires when the Claude list claims the session, which for a shell it never does. Placeholders now claim last. The card was also hard-coded toCLAUDE_KINDunder a comment reading "this app spawns Claude and nothing else (INV-7)" — true when written, false the moment a second command shape existed, and it is the second time that exact lesson has been paid for. -
A shell in a tmux session is a husk unless this app marked it. The terminal feature and
tmux-resurrect's leftovers are the same pane to anything that looks at the pane: azshsitting at a prompt in a session somebody named.is_live_agentrefuses those, and the only thing that separates the two is the@agent_commanderoptionspawn::terminal_argvwrites on the session as it creates it. It rides on the existinglist-panesformat string, so it costs no round trip — but it is also the kind of field that reads as decoration in a diff. Remove it and the terminal a user just opened never appears in the fleet, with nothing failing. -
Registry.enrich()is a blind shallow merge with two callers writing different field sets, andundefinedoverwrites. That is load-bearing for goal-clear and a trap for any new patch producer. A newAgentPatchfield has to be named in four places or it silently does nothing:applyandis_emptyinsources.rs, theoverwrite_named_fields!list intmux_source.rs, andcard_fieldsplusdiffersinenrich.rs.usagewas the last one added and touched all four.descriptionwas the next one added and touched three, which cost an afternoon. It was produced correctly,applywould have written it, andmerge_patchdid not list it — so it was dropped in between, every card showed nothing, and the whole suite stayed green, because there is no compiler and no test that reads a field's journey end to end. The half hour that went into proving the parser was right was spent on the one component that had never been wrong.Only
card_fieldsis safe by construction: it is a struct literal with no..rest, so a new field is a compile error there. The other three are now held bysources::tests::every_patch_field_is_named_by_the_three_functions_that_carry_it, which parses the struct and fails naming the field and the function that forgot it. Verified the way a guard has to be — by deletingdescriptionfrommerge_patchand watching it go red. -
A new wire type has to be named in
wire::render_once()orgen:typesemits nothing. ts-rs exports from a root list, not from the#[derive]:types.rs'srender_oncecallsexport_allon each root, and a type nobody calls it for is simply absent fromsrc/shared/wire.ts. The failure is silent in the worst way —npm run gen:typessucceeds, prints nothing, and leaveswire.tsbyte-identical, so it reads as "already up to date" rather than "never written".PictureResponsewas the last one added; the minute spent re-reading the derive was spent on the half that was right. -
tokensis output tokens only, accumulated per tail from a bounded backfill — not the session's spend, despite being presented as cost and used as a sort key. -
A half-open socket kept polling for a browser that was gone. A phone asleep behind Tailscale held a transcript tail and a share of a pane poller. Fixed by the heartbeat — then un-fixed by the port, which dropped it along with
src/server/while three documents went on describing it, and fixed again in Rust as an applicationping/pongraced against the socket read. The general rule is that nothing polls what nobody is watching; the lesson the second round added is that a property with no test carrying its number does not survive a rewrite, however well it is written up. -
Widths break in the empty state, not the full one. A 300-character search term echoed verbatim into "No agent matches …" forced the document to 2175px. The review agents found that one, plus an xterm use-after-dispose crash on every full-screen toggle and a sort that ranked unknown values as smallest.
-
A control whose success cannot be observed must not claim it failed. The permission-mode button was reported broken three times across two rewrites, and the key was never at fault:
tmux send-keys BTabemits\033[Z, and three presses walk a live sessionauto→planexactly as the keyboard would. Claude Code writes itspermission-moderecord at the end of a turn, so an agent at its prompt — the one usually being switched — reports nothing back, and one that has not taken a turn has no transcript to report from. Verifying it meant a 2.5s window with the button dead, thenunverified, then the old mode still on the label: three signals of failure about something that had worked. It now sends the key and says so. Before building verification onto a control, check that the CLI writes anything down when the thing happens. -
A control action that carries no value still goes through
readJson. Mode, clear and compact take no argument, so the browser sends no body at all — andJSON.parse('')throws. Every unit test passed while pressing the mode button reported "that did not take effect: Unexpected end of JSON input" about a control the server had never called. The Rust port keeps both halves of that lesson:routes::inv8_mode_and_clear_and_compact_carry_no_bodydrives the real HTTP handler with an empty body, becausecontrol::testscalls the action functions directly and structurally cannot catch this. -
/clearreplaces the session; it does not edit it. Claude Code opens a fresh transcript under a new session id and rewrites~/.claude/sessions/<pid>.json. Anything holding the old id — a URL, a socket focus, a mock fixture, an e2e test — is stale the moment it lands. That is also why one fixture (AGENT.clearable) is reserved for the clear spec: every e2e project shares one mock server, and clearing a fixture another test uses deletes that test's agent out from under it. -
npm testis flaky under load, andschemeis not the only one.test/scheme.test.ts(a 5s timeout around spawningpython3) fails when the machine is busy and passes on a quiet one.test/ui/token.test.tsx("keeps it when opening an agent" / "…closing one again") does the same: measured on a cleanmainwith nothing else changed, it failed 3 runs out of 6 back-to-back and passed 8 of 8 when run on its own. Both are load, not regressions — but check that way round before believing either, because a red run that names only these is the cheapest kind of false alarm to chase. A red Stop hook naming only that is worth re-running before believing. The INV-4 tail-count flake that used to sit beside it went away with the port:enrich.rs's cadence re-arms after the work instead of on a wall clock, so a slow pass no longer drops a tick. -
WebKit times out under a loaded machine, and it is not always the same spec. Measured on 2026-09-06 at load average 8.7: a full run failed
responsive.spec.ts"every control on screen can be hit and announced" onphone-safariandtablet.spec.ts"both columns are visible at once" ontablet-safari, both withTest timeout of 30000msonpage.gotoor on tearing the context down, twice in a row — and the same two failed with the working tree stashed back to the released tag, which is what proves it is the machine rather than the change. Both passed 2 of 2 run alone. The signature to match is: WebKit only, a 30s timeout rather than a failed assertion, and near the end of a project's run. Stash and re-run the baseline before believing a WebKit failure belongs to your diff. -
theme.spec.tson the two WebKit projects flakes under a full-suite run. Two of its tests (picking a scheme repaints the documentandthe scheme and the theme both survive a reload) failed withpage.goto: Test timeout of 30000ms exceededontablet-safari/phone-safariin two consecutive full runs, and passed 28 of 28 when the spec was run alone on those projects. It is the WebKit navigation stalling under load, not the app: a red full run naming only those is worth a rerun of that spec before it is believed. -
A test that scrolls and then waits is a race, not a wait.
fades.spec.tsscrolled the detail pane to its end and waited for the fade to lift. That pane's content is still arriving — the conversation lands over the socket a beat after the answer card — so a re-render between the scroll and the read put the box back at the top, with the fade correctly reporting that there was more below. It failed in two full runs and passed 7 of 7 on its own, which is the signature: the assertion held the content still, not the thing it was about. The scroll now happens inside the poll. -
The e2e
/clearfollow test flakes on slow CI runners.control.spec.ts"INV-8 follows the agent to the session it is now running" failed both attempts on one GitHub runner and passed on rerun with nothing changed. One red occurrence of exactly that test is worth a rerun before it is believed — and worth root-causing if it ever fails twice in a row on different runs. -
The Mac bundle's path constraints went from four to one, and the one that survived changed shape. The Node bundle needed
dist/webbesidedist/server,scripts/statusline-bridge.mjstwo levels up,dist/sharedas a sibling, and apackage.jsoncarrying"type": "module"— without that last one Node read the ESM output as CommonJS and died before printing a character, a do-nothing binary wearing a Dock icon. Three of those were Node's and left with it. What remains is the web root: the launcher passes--web-rootexplicitly, andResources/webis laid out as a sibling ofResources/binsodefault_web_root()'s own../webfallback lands in the same place for anyone running the staged binary by hand.--helpreturns before anything is served, so the smoke test cannot prove this one —build-mac-app.pychecks forweb/index.htmldirectly instead. -
npm run e2ereuses a running server, so a Rust change is invisible to it.playwright.config.tssetsreuseExistingServer: !process.env.CI. ThewebServercommand builds the bundle and the binary — but only when Playwright actually starts one, and if something is already listening on 4599 it attaches to that instead and builds nothing. The bundle still updates, because the server readsdist/webfrom disk on every request; the binary does not, so every mock fixture, route and wire change is silently the old one.This wasted several bisect steps: a fixture was removed, the suite still failed, and the DOM still contained the thing that had been deleted. The tell is exactly that — a failure that mentions something no longer in the source. Before trusting any e2e run that turns on server behaviour:
lsof -nP -iTCP:4599 -sTCP:LISTEN -t | xargs -r kill lsof -nP -iTCP:4598 -sTCP:LISTEN -t | xargs -r killIt is the same lesson as "a running server outlives the bundle it serves", one layer up, and it is worse there because the harness looks like it owns its own server.
-
A 16px icon added to a topbar chip broke four terminal tests, and the connection is four steps long. The hand went on the "N need you" filter chip;
attach.spec.tsthen failed everyterm-wrapclick with "element is not stable", for the full 30s, on desktop. The chain: the topbar is a wrapping flex row, an SVG in a chip moves that row's height,mainresizes, andPaneTermre-fits the terminal on its parent's resize — which changes layout again.term-wrap's box therefore never produced the two identical animation frames Playwright's click stability check waits for.Three things worth keeping from finding it. It reproduced at load average 6, so "WebKit under load" was the wrong first guess and cost two runs. The baseline is what settled it —
git stash -u, rebuild, run the one spec: 9 passed in 13.7s against 2.2 minutes of timeouts, which is not a flake signature. And the bisect had to go through five candidate groups (the card, the server, transport,Message, and finallyApp) because the failing surface and the changed file share nothing but a layout ancestor. A continuously repainting pane makes any layout change anywhere a candidate; the terminal is the most sensitive thing on the page and it is never the thing you changed. -
codesignruns before the smoke test, not after. The launched thing is a Mach-O now rather than a script, and on Apple silicon an invalid signature is a kernel kill rather than a warning. Signing first means the smoke test runs the exact bytes the user will. It stays non-fatal, so a benign codesign failure does not stop a build while a lethal one cannot slip past it. -
The bundle ships no
statusline-bridge.mjs, deliberately. A bridge path written into~/.claude/settings.jsonthat points inside a.appbreaks the next time the app is replaced, so--install-statuslinebelongs to the npm package. Run from inside the bundle it now reports that it cannot find the bridge script rather than writing a path that will rot. -
A bundled app keeps serving the code it started with. The
.appcarries its own copy of the binary, so reinstalling replaces the bundle and changes nothing about the server already running — and the launcher's own already-running check then finds a healthy server and just opens the browser, so the update lands with no effect and no message. The launcher compares the bundled binary's mtime against the pid file's and restarts a server it started itself. Timestamps rather than versions: a version only moves on a release, so rebuilding at the same one all afternoon would defeat a version check in exactly the case that happens most. It replaces only a server whose pid file it wrote and whose command line names this bundle, so a copy you started from a clone is opened, never killed. -
actions/upload-artifactdoes not preserve the executable bit. It zips what it is given, and the zip carries no mode, so a binary that left the build job at 0755 comes back fromdownload-artifactat 0644 andexecFileSyncraisesEACCES. Nothing in the release job notices — the tarball is well-formed, the publish succeeds, and the first person to runnpxgets the failure.npm-publish.ymlthereforechmod 755es each binary as it lays it out and then asserts the mode came back 755 before publishing. Both lines are load-bearing rather than defensive. This is the one entry here that has not bitten yet; it is on the list because its failure mode is the do-nothing install for the third time, and because achmodwith no comment on it is exactly the line someone tidies away. -
npm pack --dry-run --jsondoes not have a stable shape, and it broke the first release it was guarding. The publish job asserted the four binaries were really in the tarball by parsing that JSON as an array — verified locally against npm 11.17, where it is one. The workflow's ownnpm install -g npm@lateststep then handed the runner a newer npm whose output it could not read, sopacked[0].filesthrew and the job died inside the guard, with the package itself perfectly fine. It failed safe, which is the one good thing about it. The guard now runsnpm packfor real and reads the archive withtar -tzvf: the tarball is the artifact being uploaded and its format does not move between npm releases, and the listing carries the executable bit directly. A check that parses another tool's optional output format is a check that can fail for reasons unrelated to what it is checking. -
Piping a test run through
taileats the verdict.npm run e2e | tailreports tail's exit code, not Playwright's, and the failure list scrolls out of the kept lines — a 92-failure run once read as "141 passed" that way. Redirect to a file and check the exit code, never pipe a gate. Backgrounding one has the same shape:cmd > log &followed byecho $?reports the backgrounding, not the run. Put theechoinside the backgrounded shell, or read the verdict out of the log — a 5-failure run read as "exit 0" that way. -
A redirect can outlive the reason for it. Opening a blocked agent jumped straight to the Attach tab, on the stated grounds that it was "the only tab that can answer its dialog". Once the Chat tab could answer one (INV-16), that premise was false and the redirect was actively carrying users away from the better surface — but it kept working, so nothing failed. Its own comment is what gave it away. When you add a capability, grep for the comments that assert it is impossible.
-
openAgentwaits for a first message, and not every fixture has one.e2e/helpers.tsblocks ongetByTestId('message'), so it can never open the two fixtures that were never prompted — includingmock-waiting, which is exactly the one a blocked-agent test wants. That is a real shape rather than a quirk: an agent asks for permission on its first tool call, before it has said anything. Navigate directly for those. -
A browser gate that opens a viewport without
hasTouchmeasures a device nobody is holding. Everypointer: coarserule is off while it runs, so a 44px touch floor that exists only behind that query is invisible to the audit and to the person reading its clean output.Chat.module.cssdocuments this at.sendModeOption, which carries a width query and a pointer query for exactly this reason — and it bit again anyway atMessage.module.css's.toggle, which had only the pointer query and sat at 24px against WCAG 2.5.8's 24 and the 44 a finger wants. A touch floor needs both queries: the width cut for the audit and for a phone, the pointer query for a 1194px landscape tablet that is too wide for the cut and still a screen people tap. The same blindness ran the other way inaudit-workspace.mjs, which measured the tablet at 80.9% with a fine pointer and 79.5% with the coarse one it actually has. -
Playwright's default click waits for the element to be stable, and beside a live capture that wait can never end.
page.clickrequires two consecutive animation frames with an identical bounding box before it acts. The Attach tab repaints continuously, so on a loaded machine — measured here at load average 14.7 — the pair never arrives, the click retries to its 30s timeout, and the script dies before printing its report. That is the dangerous part:audit:a11yfailed twice in a row and the output contained no findings at all, which reads exactly like a clean run to anyone skimming. The three audit scripts now locate, wait for visible, and click withforce, which keeps every meaningful precondition and drops only the one that cannot hold next to a capture. If an audit ever exits non-zero with no===== AUDIT =====line, it did not run. -
A guard is not verified by watching it pass.
test/icons.test.tssweeps for icon buttons that lost their accessible name, and it shipped toothless: it stripped JSX tags with/<[^>]*>/, which breaks on the>insideonClick={() => f()}, so leftover code read as "this button has text" and every such button looked labelled. Deleting a realaria-labeldid not fail it. A new guard has to be proved the other way round — break the thing it guards, watch it go red, put it back — and anything parsing JSX with a regex needs a brace-aware scan rather than a character class. -
A
vi.fn(() => …)cannot benew-ed. StubbingglobalThis.WebSocketwith an arrow function makesnew WebSocket(url)throw "not a constructor" before the body runs, so the mock records zero calls while the code under test still sees a throw — which looks exactly like the code never calling it. Usevi.fn(function () { … })for anything constructed.
Shipping it to npm
One package carries every binary:
dist/bin/darwin-arm64/agent-commander
dist/bin/darwin-x64/agent-commander
dist/bin/linux-x64/agent-commander
dist/bin/linux-arm64/agent-commander
dist/web/… the Vite bundle, unchanged
.github/workflows/npm-publish.yml builds one binary per matrix job and a final
job collects the four, marks them executable, and publishes once. Four binaries
is ~7.5 MB, and that is the deliberate trade: the optionalDependencies matrix
esbuild and Biome use would ship a quarter of the bytes but has a window during a
release where the platform package a user resolves is not on the registry yet,
and what they get is an install with no server in it. This is a global CLI
installed on purpose, not a transitive dependency; the bytes are cheap and that
window is not. One package also means one trusted-publisher configuration and
nothing to reconcile when a matrix job fails.
The directory name is exactly ${process.platform}-${process.arch}. Not a
convention that happens to line up — scripts/launch.mjs interpolates those two
values and looks there, so there is no mapping table to maintain and no way for
the launcher's names and the release job's names to drift apart. Rust's target
triples (aarch64-apple-darwin and friends) appear only inside the workflow,
where the translation to a directory name is written down once. Adding a target
means adding a matrix row; nothing in the resolver changes.
Windows is deliberately not a target, and not a gap to be closed later. The
Attach tab is tmux capture-pane and send-keys, and the fleet is read out of
~/.claude/sessions/<pid>.json. A Windows build would install cleanly, start,
and command nothing. Adding one is not a matrix row, it is a second
implementation of INV-1.
npm run build:server still writes rust/target/release, and must keep doing
so. Eight things read the binary from that path: npm run dev, mock,
start and serve; playwright.config.ts's webServer;
scripts/build-mac-app.py; scripts/dist-bin.sh; and launch.mjs's own second
candidate, which is what makes a checkout runnable without a publish ever having
happened. dist/bin is a layout assembled out of four separate runners' output
— not something any one machine's build produces in full. Repointing
build:server at it would break all eight to save one copy.
npm run build:dist-bin (scripts/dist-bin.sh) assembles that layout from
whatever the local rust/target happens to hold, which is how a packaging
change gets verified without a tag: build, run it, npm pack, read the tarball.
Note the two shapes it handles — a cross-compile lands in
rust/target/<triple>/release while the host's own build lands in
rust/target/release with no triple at all.
Delete dist/bin when you are done with it. It is the launcher's first
candidate, so one left behind in a checkout shadows rust/target/release: from
then on npm run build:server changes the binary the e2e harness and every
npm run server use, while the one bin runs stays whatever you assembled that
afternoon, and nothing says so. dist/ is gitignored, so git will not remind you — and files
ships dist wholesale, which is the same trap the .gitignore comment on
build/ already names.
There is no postinstall, nothing is downloaded at install time, and a Rust
toolchain is needed only to work on the server. Never verify by publishing a
throwaway version: npm versions are immutable, so the only way to withdraw a bad
one is to burn the next number.
npm-publish.yml is commented at length and those comments are the reference,
not this section. Two of them decide things you would otherwise change by
accident: the Linux legs are pinned to ubuntu-22.04 rather than
ubuntu-latest because a glibc-linked binary runs on the glibc it was built
against or newer and never older, so bumping the runner raises the floor under
every user at once — on their machine, after publish, where CI cannot see it.
And darwin-x64 is the one target nothing in CI executes, because an arm64
macOS runner has no guaranteed Rosetta; it is built and shipped unrun, and the
publish job's mode and tarball checks are what stand in for a smoke test there.
The macOS bundle
npm run build # the Vite bundle, then cargo build --release
python3 scripts/build-mac-app.py # --out defaults to build/
One self-contained binary in Contents/Resources/bin plus the Vite bundle in
Contents/Resources/web; Contents/MacOS/agent-commander is launcher.sh,
which probes the port, detaches the server, opens a browser and exits. It
detaches rather than execs because launchd kills the job's process group —
the Node bundle used node's own spawn() for that and there is no interpreter
left to borrow, so it forks through /usr/bin/perl with a logged
plain-background fallback.
build-mac-app.py runs the staged binary with --help before promoting the
bundle. That check is why the do-nothing bundles above were caught rather than
shipped, and it is worth keeping whatever else changes.
It is not covered by any gate. npm test does not touch it, no e2e test
builds it, and .claude/gates.json reaches it only through the scripts/
watch entry, which triggers the three cheap gates rather than a bundle build.
Build it by hand after changing anything it stages.
Review agents
Point harness:qa-bar-raiser and harness:ux-bar-raiser at a --mock server on
4400 and never at 4317. Both review only — they never edit code — and both are told to say
"nothing found" plainly rather than pad a list. docs/HANDBOOK.md §"Review
agents" lists what they have caught.
The three documents worth reading in full
ARCHITECTURE.md— the module graph, the five planes, what is pushed versus polled, and §"Where it is fragile", which is ordered by how quietly each thing fails. Trim an entry when it is fixed; the record of trimmed ones is §"Fixed since this list was written". §"How it is checked" is the gate design.INVARIANTS.md— INV-1 … INV-18, each with the tests that prove it.CLAUDE.mdimports it alongside this file, so it is in context for every session here without being asked for.SPEC.md— what the app is supposed to do, requirement by requirement, in more detail and more fluently than the invariants allow. The invariants are the subset of it whose violation is a defect; where the two disagree the invariant wins. It has no gate, so it is held down by citing a test for every requirement it makes.
README.md is the one-page version for the person installing the app, and
docs/HANDBOOK.md is the same at length. These three are for the person
changing it, and CONTRIBUTING.md is the short route from a clone to a pull
request.
TODO.md is queued work, written to be executed cold by whoever picks it up:
what the change is, the call sites as they stood, and what "done" is checked
against. Take an item or leave it, but read it before proposing one of your own
— it also records what was considered and deliberately rejected, so a rejected
idea does not get re-proposed as a new one.
Commits
- Explain why, not just what. The body is where the reasoning goes.
- Never add a
Co-Authored-By: Claudeor any AI-attribution trailer. - Commit or push only when asked.
.cleancode.json turns magic-number down to a warning, and only that
rule. The pre-commit hook relaxes its thresholds for test code, but it decides
what is test code from the path — so test/ui/*.test.tsx is relaxed and a
#[cfg(test)] mod tests inside rust/src/usage.rs can never be, because Rust
does not put its tests in a test path. Of the 28 findings that blocked the
0.17.0 release, 16 were exactly that: assertion literals a line below the
fixture they round-trip, where lifting the number into a shared constant would
stop the test from catching a transcription error rather than help it. The rest
were structural — a byte offset into a literal buffer, splitn(4, '/') for a
URL path, the 0.01 named by the '<$0.01' beside it. Every other rule still
blocks, and the hook's secret scan cannot be disarmed from in here at all.
