Imported from shoedog/a2acp (
AGENTS.md). Install upstream withnpx skills add shoedog/a2acp. Copyright stays with the author.
Using a2a-bridge (agent quickstart)
Start here: read the
a2a-bridge-operatorskill before running or diagnosing an agent workflow. Checkdocs/compatibility.mdfor tested versions and incident dispositions, and use the checked-in compatibility and dogfooding routing reference when selecting a model or review lens. Do not infer host support from a container result or vice versa.
a2a-bridge is an A2A↔ACP bridge and a multi-agent workflow runner. You can use it as a tool to run
clean-room design, code/spec/plan review, or autonomous implement passes against any repo —
each step driven by real coding agents (codex, claude, kiro, …) over the Agent Client Protocol.
If you were sent here to "run a workflow / review / design through the bridge," this file is all you need.
Do NOT read bin/a2a-bridge/src/*.rs to find the invocation — it's below, and every subcommand has
--help.
0. Where configs, prompts, and workflows live
Keep this repo's examples/ and prompts/ for generic bridge examples. Codebase-specific workflow material
belongs with the codebase that owns it, not in a2a-bridge; for example, Prism/slicing workflows should live
under that repo, such as tools/a2a-bridge/configs/ and tools/a2a-bridge/prompts/. Disposable one-off
configs/prompts/workflows should live under /tmp (or /private/tmp on macOS) or another scratch directory.
Before serving or handing a config to another agent, run:
a2a-bridge validate --config /path/to/a2a-bridge.toml
Before committing local changes in this repo, run the repository hygiene guard:
cargo run -p a2a-bridge -- validate --repo-hygiene
Use --examples-policy deny in cleanup/CI gates when you want to reject project-specific workflow material
under an examples/ directory; pass the strings that identify that project with repeated
--project-marker flags, for example --project-marker code/slicing --project-marker prism-mcp.
a2a-bridge mcp is an external-controller surface. Never point a bridge-managed agent's
[[agents.mcp]] entry back at a2a-bridge mcp. Direct loopback configurations fail validation, every MCP
spec delivered to a managed agent receives the reserved A2A_BRIDGE_MCP_CALL_DEPTH=1 marker, and a marked
bridge MCP invocation refuses before config, store, or coordinator work. Do not set that reserved variable
in agent MCP configuration. This is a safety guard against accidental recursion, not a security boundary
against a deliberately hostile wrapper that removes inherited environment variables.
1. Build / install
cargo build --release --bin a2a-bridge # → target/release/a2a-bridge
# or: cargo install --path bin/a2a-bridge # → ~/.cargo/bin/a2a-bridge
a2a-bridge help # top-level usage; <subcmd> --help for details
2. Run a workflow against ANY repo
a2a-bridge run-workflow <id> \
--input brief.md \ # the problem statement / material to act on
--session-cwd /path/to/target-repo \ # the repo the agents read/work in (NOT the launch cwd)
--config examples/a2a-bridge.multi-agent.toml \
--out result.md # omit to print the terminal node to stdout
- The terminal workflow node's output is what you get (stdout or
--out). Runs offline. --session-cwdis the per-request cwd (ADR-0014). Without it, agents run in the launch cwd, not your target repo — a common mistake.
Built-in workflow <id>s (defined in examples/a2a-bridge.containerized.toml and
examples/a2a-bridge.multi-agent.toml):
design (2 clean-room architect lenses → synth), code-review (one Sol/xhigh read-only risk-triaged pass),
spec-review, plan-review.
A workflow is just [[workflows]] + [[workflows.nodes]] in the config — copy one to make a variant
(e.g. a codex-only design).
3. Implement a task in a repo (clone → edit → verify → review → commit)
a2a-bridge task-spec template implement > task.md # scaffold a typed task-spec; edit it
a2a-bridge implement --input task.md \
--repo /path/to/target-repo \
--config examples/a2a-bridge.containerized.toml \
[--depth auto|light|standard|thorough] # override the auto review depth (optional)
Clones the repo into a quarantine under allowed_cwd_root, runs the warm containerized impl agent
(edit + fix turns share ONE container + session), build/test-verifies, reviews the diff, and hands off a
branch for you to merge. The default impl agent is codex (gpt-5.5, effort=high).
Convention: the containerized verify step above already runs with CARGO_INCREMENTAL=0 (set in each tracked
config's rust verify_env, plus CI). Do the same for any other one-shot full-suite/aggregate-verifier run
(e.g. a manual cargo test --workspace pass) — a fresh invocation never reuses incremental artifacts, and
leaving it on only bloats the cache (measured: 44% of a 15.77 GiB verify cache was incremental artifacts).
Interactive/warm dev sessions are unaffected.
The shipped review-the-diff default is one host-side gpt-5.6-sol/xhigh hard-read-only review at every
size tier. It completes correctness findings first, then revisits every WRONG/SMELL for real-world trigger
conditions, likelihood, exposure, bounded repair cost, and blocker/defer value. Auto-sizing and --depth remain
available for explicitly configured alternate tier workflows (persisted across --resume).
Land it (merge, ADR-0027). Integrate an Approved run's commit into its source repo, re-authored to
you (the operator), without touching your working checkout:
a2a-bridge merge <id> --onto main # land run <id> onto `main` (fast-forward off its base_commit)
a2a-bridge implement --input task.md --repo … --merge --onto main # implement + auto-merge when Approved
a2a-bridge merge <id> --onto main --integrate-current # inspected parallel sibling; compose onto current main
merge re-authors the clone's commit via git commit-tree and lands it with
git push --force-with-lease=refs/heads/<target>:<base_commit> (the lease IS the concurrency CAS — one of N
concurrent merges wins, the rest get a stale-lease refusal). Operator identity comes from the source repo's
git config user.name/email (or a [merge] author_name/author_email override). Exit codes: 0 merged ·
1 usage/preflight · 2 (--merge) run not Approved · 3 Approved but couldn't land (target moved, checked
out, diverged, or conflicted). Default merge remains exact-base Mode A. For a reviewed sibling from a shared
immutable base, --integrate-current three-way composes its delta onto the fetched current target and leases
that exact target commit; conflict or non-ancestor history refuses with the clone retained. Resume and merge use
one clone-local operation lock. Caveat: a source repo with
receive.denyCurrentBranch=updateInstead/ignore is out of scope (the default refuse is the no-touch backstop).
Parallel implementor flight (ADR-0040). Freeze one base SHA, write independently testable task specs, and
assign every changed path/seam to one implementor; reserve shared manifests, roadmaps, generated files, and
cross-cutting cleanup for an integration task. Start all siblings with the same --base-ref <sha> and one shared
config, but do not give siblings --merge. Inspect every terminal result, then integrate Approved siblings one at
a time in dependency order with merge --integrate-current. A conflict stops the flight for an ownership/fix
decision; it never triggers an automatic agent retry. Run the aggregate full suite and review the combined
<base>..<target> diff after the final landing. The reliability roadmap remains the sole program cursor.
4. Serve (A2A server)
a2a-bridge init --agents codex,claude # scaffold ./a2a-bridge.toml + prompts
a2a-bridge serve --config ./a2a-bridge.toml
serve advertises each agent's available models/effort/modes on the Agent Card
(agent-models extension, probed at startup + refreshed on SIGHUP).
4b. Discover model/effort/mode values
a2a-bridge models --config ./a2a-bridge.toml # table: each agent's advertised models/effort/modes
a2a-bridge models --config … --agent codex --json # one agent, JSON (caps or explicit failure)
Probes live and degrades per-agent. Successful JSON values use the Agent Card agent-models capability
shape. A failed probe is retained as {available:false,failure:{agent,strategy,phase,...}}; an explicitly
requested failed agent exits nonzero after printing that machine-readable record, while an all-agent probe
keeps partial-success exit behavior. Failure detail is returned in place: text includes the stable phase,
category, and deepest bounded redacted error; JSON exposes those as failure.phase, failure.category, and
failure.error, plus failure.diagnostic when the ACP boundary supplied a typed diagnostic. There is no
implicit external “see logs” destination. Pass any listed value to the per-request override
(message.metadata a2a-bridge.{model,effort,mode}) or an agent's config default.
4c. Run one explicit live smoke (billable)
Only after the operator explicitly authorizes a billable turn, use the candidate release binary for one
fixed, bounded PONG probe:
evidence_dir="$(mktemp -d /private/tmp/a2a-bridge-smoke.XXXXXX)"
chmod 700 "$evidence_dir"
cargo build --release --bin a2a-bridge
./target/release/a2a-bridge smoke \
--agent codex \
--config /absolute/path/to/a2a-bridge.toml \
--model gpt-5.6-sol --effort xhigh \
--session-cwd /absolute/path/to/trusted-repo \
--timeout-secs 120 \
--acknowledge-billable \
--out "$evidence_dir/codex-host-smoke.json"
The command sends only Reply exactly PONG. Do not use tools., resolves/configures/prompts once, never
retries or falls back, and requires both exact PONG and a successful terminal event. Missing billing
acknowledgement and malformed options refuse before config/registry/spawn work. Once argument and output
preflight passes, an acknowledged attempt writes its versioned artifact before returning nonzero.
Without --out, stdout is JSON only; human direction goes to stderr. Do not pass
--include-redacted-stderr unless bounded best-effort-redacted process text is specifically required.
An explicit output path must not already exist. On Unix, it is created owner-only as 0600 before agent
resolution or spawn; an existing file/link or failure to apply that restriction is a pre-attempt refusal.
Run validate, doctor --json, and models --agent <id> --json first. Claude smoke refuses before adapter
spawn when bounded OAuth metadata is expired or has less than 16 minutes of runway; syncing an isolated
credential copy does not refresh an expired host login. When CLAUDE_CONFIG_DIR is present for a host
Claude entry, it must be a non-empty absolute path so doctor and every possible child cwd select the same
.credentials.json; unset uses $HOME/.claude/.credentials.json. The one smoke deadline begins before
provenance and orphan recovery, so those phases cannot consume the runway and then receive a fresh timeout;
one deadline-first primitive refuses without polling resolution, configure, prompt, or drain when time is
already exhausted. A stage is counted only after its future receives a poll; an unpolled prompt refusal
records zero prompt calls and false prompt-acceptance evidence. Truthy
CLAUDE_CODE_USE_BEDROCK, CLAUDE_CODE_USE_VERTEX, CLAUDE_CODE_USE_FOUNDRY,
CLAUDE_CODE_USE_ANTHROPIC_AWS, or CLAUDE_CODE_USE_MANTLE selects external provider auth and therefore
skips first-party file OAuth on host entries; false-like or unknown values do not, and ambient host flags
never bypass a mounted reader credential.
Never use a stale installed binary
for compatibility evidence, and never automatically rerun a failed or timed-out smoke: the first prompt may
have been accepted. Do not update docs/compatibility.md until the release-mode artifact records the exact
lane that actually ran.
4d. Validate or run the compatibility matrix
The checked-in compatibility manifest is non-billable to validate:
a2a-bridge compatibility validate --manifest compatibility/manifest.toml
Running cases is potentially billable and therefore requires an explicit lane/case selection, the environment owner, an acknowledgement, and a new aggregate output path:
a2a-bridge compatibility run \
--manifest compatibility/manifest.toml \
--lane pinned \
--environment-owner <manifest-owner> \
--acknowledge-billable \
--out /private/tmp/compatibility-aggregate.json
The runner canonicalizes and descriptor-pins the aggregate parent, creates aggregate and scratch entries
relative to that retained descriptor, rechecks its identity during creation, and refuses normal or bare
Git repository ancestors. The new mode-0600 output immediately contains a blocking setup-incomplete
aggregate, so later scratch/staging failure remains valid evidence instead of an empty file; keep
compatibility evidence in disposable operator-owned storage. Manifest prerequisites use
structured entries: { name = "PATH" } means presence-only, while
{ name = "A2A_BRIDGE_ALLOW_FABLE", one_of = ["1", "true"] } binds accepted non-secret values.
Pinned adapter/CLI values require one complete semantic <package>=<version>. Remote API rows require
dedicated provider, api, and api_version component identities; a generic execution row is not a pin.
An alias-shaped model ID may be an exact advertised raw ID, so the runner also requires the successful
effective model to equal the requested pin and blocks a fallback alias resolution as drift.
Each eligible case invokes one bounded, privately staged snapshot of the exact candidate binary's
fixed-PONG smoke once. The aggregate records its SHA-256 and byte length; the runner refuses digest
drift, publishes the staged inode owner-executable but non-writable as mode 0500, executes the
verified file object instead of reopening its name, and accesses child smoke
artifacts relative to the retained scratch descriptor. After hashing it rechecks cancellation and the
full declared timeout headroom immediately before spawn. On Linux, the staged child closes its inherited
candidate descriptor after exec and its scratch descriptor after opening the artifact, before ACP
descendants. A pinned config's exact SHA-256 is an admission gate before provider spawn. Container pins
require exact non-secret adapter/CLI labels from the configured immutable image, and the Fable reader
also binds exactly one minimal host-mounted settings file by SHA-256; missing, unreadable, or ambiguous
duplicate settings destinations cannot green a support case.
There is no retry, provider fallback, implicit all-case selection, baseline update, or production-config
mutation. A case does not start unless its declared token and observable-cost caps fit the remaining
total headroom. Negative/non-finite cost observations fail explicitly and remain sticky across later
usage snapshots. Comparison retains per-case
execution/error/not-run/budget state and aggregate success/cancellation/budget state while excluding
variable usage quantities. The checked-in manifest has reviewed R3b pins; its baseline is promoted only
from separately authorized exact-candidate evidence. Read
docs/compatibility.md and the current
reliability roadmap before spending a live turn.
Floating-current canaries require two independent authorizations. First validate the request without effects, then authorize its exact registry/image effects and a new private output directory outside every Git repository:
./target/release/a2a-bridge compatibility validate \
--recipes compatibility/floating-current.toml
evidence_root="$(mktemp -d /private/tmp/a2a-bridge-floating.XXXXXX)"
chmod 700 "$evidence_root"
./target/release/a2a-bridge compatibility resolve \
--recipes compatibility/floating-current.toml \
--case <exact-floating-case-id> \
--environment-owner <exact-owner-id> \
--runtime docker \
--acknowledge-resolution-effects \
--out "$evidence_root/resolution"
The recipe's selectors are requests, not compatibility evidence. Resolution may use the npm registry,
runtime cache, and one unique disposable image tag, but it starts no adapter/provider session, calls no
models, copies no credentials, replaces no shared tag, and grants no billing permission. The npm
subprocess may create only the exact lock through the bridge's fixed npmjs CONNECT proxy; it never receives
tree-write authority.
The bridge downloads the exact integrity-bound npmjs HTTPS archives, requires matching package identity and
present declared bin targets, and preflights paths/types in a case-insensitive portable ASCII namespace. It
raw-preflights every GNU long-name/long-link and local/global PAX metadata record against a 1 MiB cap before
tar preprocessing in both planning and materialization, accounts PAX-effective file sizes, binds symlink
targets to the exact spelling of portable-equivalent planned paths, reserves the entire aggregate entry/byte
budget before the first package entry write, and materializes the private tree descriptor-relatively.
Inspect the complete resolution.json, validate and doctor its generated configs, then obtain separate
authorization
for the exact resolution id, unchanged candidate binary, selected cases, owner, and budget:
./target/release/a2a-bridge compatibility run \
--resolution "$evidence_root/resolution/resolution.json" \
--all-resolved \
--environment-owner <exact-owner-id> \
--acknowledge-billable \
--out "$evidence_root/floating-aggregate.json"
./target/release/a2a-bridge compatibility compare \
--current "$evidence_root/floating-aggregate.json" \
--mode floating-to-pinned
Use --all and --all-resolved only as explicit authorizations. The run revalidates every bound artifact
immediately before provider spawn and captures the bounded catalog from that same one-prompt session.
candidate_pass, candidate_fail, and candidate_unknown are advisory canary outcomes; none promotes or
rewrites production pins, baselines, configs, support docs, or the running operator. Retain the private
bundle and unique tag until operator-reviewed cleanup proves that no running container uses them. Floating
comparison rejects a baseline whose pinned-manifest identity differs from the resolution-bound production
manifest.
4e. Plan an explicit host verification after classified container degradation
Current slice status, review evidence, sequencing, and handoff are owned solely by
docs/reliability-execution-roadmap.md. This file defines the
stable operator behavior and must not duplicate changing candidate hashes or gate totals.
Only a complete failed smoke schema-v2 artifact can be evaluated. The source config must still be the
same canonical regular file with the same SHA-256, its configured source agent must still be a read-only
container using the same canonical mount, and the target must be an unsandboxed ACP entry explicitly
marked host_fallback_eligible = true:
./target/release/a2a-bridge fallback-plan \
--from /absolute/path/to/failed-container-smoke.json \
--host-agent trusted-host-review \
--config /absolute/path/to/a2a-bridge.toml \
--trusted-session-cwd /absolute/path/to/exact-owned-repo \
--confirm-trusted-own-repo-read-only \
> /private/tmp/fallback-plan.json
The command is local and non-billable. It accepts only a pinned, bounded regular-file smoke-v2 artifact;
hand-assembled task envelopes and historical smoke-v1 artifacts are not trusted fallback evidence. An
ineligible plan contains no command. An eligible plan emits an absolute candidate-binary argv for a
distinct fixed-PONG verification smoke, bound to the current executable/config SHA-256, source-agent
marker, and the plan-time source mount's canonical path plus descriptor-derived persistent-object
fingerprint. The separately supplied trusted cwd must be an existing canonical directory, must exactly
match the artifact-reported cwd as evidence, and must remain under that mount snapshot. Only that exact
operator-selected directory enters the host smoke argv, and its own plan-time canonical value plus a
descriptor-derived persistent-object fingerprint are separate closed-set guard fields. Filesystems
without a durable object ID/handle fail closed.
fallback-plan never runs the emitted argv. Inspect the JSON and explicitly decide whether to invoke it;
the generated smoke still contains --acknowledge-billable. At action time the smoke re-reads the config
and executable and revalidates the exact cwd object, the exact source-mount object and containment, and
the target marker before any agent spawn. Same-mount symlink/sibling, mount-symlink retarget, or
inode-reuse replacement fails closed. Because the guarded target is already proven to be unsandboxed
ACP, guarded composition ignores its configured session_cwd/cwd aliases and uses the pinned
object-addressed cwd for native MCP/Kiro inputs, process redaction, and ACP session configuration. That
smoke does not call the container runtime for recovery or run-end cleanup and records the backstop as
not_needed. Never reconstruct or omit the generated guard flags by hand, and never treat a fixed
PONG as a retry/resume of the original task.
5. Inspect / clean up containers
a2a-bridge containers list --config examples/a2a-bridge.containerized.toml # this config's containers
a2a-bridge containers list --config examples/a2a-bridge.containerized.toml --all # every managed container
a2a-bridge containers reap --config examples/a2a-bridge.containerized.toml # reap DEAD (crashed) only
a2a-bridge containers reap --config … --all-dead # every owner's DEAD
a2a-bridge containers reap --config … --force a2a-rw-<owner>-<run>-0 # reap one by name (any state)
list classifies each container alive / dead / unknown by probing its run's flock lease (a free lock
⇒ the owning run crashed) and flags stale ones (no output within --older-than, default 1h). Reap is
Dead-only by default — a live concurrent run is never touched; --stale reaps idle-but-alive,
--force <name> is the only override (also how you clear legacy pre-Increment-A containers).
cwd, configs, creds, concurrency
- cwd:
run-workflow→--session-cwd;implement→ derived from--repo(it clones it).serve→ per-request via the A2A message metadata. - Configs:
examples/a2a-bridge.containerized.toml(containerized agents behind an egress lock + theimplement/verify/review blocks),examples/a2a-bridge.multi-agent.toml(host agents + the review/design workflows), ora2a-bridge init. - Creds (containerized agents): WRITABLE single-file copies in
~/.config/a2a-creds/{claude,codex}—cp ~/.codex/auth.json ~/.config/a2a-creds/codex/auth.json, likewise claude (its OAuth token expires ~hourly, so re-copy if a claude node starts failing). On macOS hosts whose Claude login is Keychain-only,deploy/containers/sync-creds.shbuilds the claude copy from a long-livedclaude setup-tokentoken (CLAUDE_CODE_OAUTH_TOKENor~/.config/a2a-creds/claude/oauth-token). Seedocs/containerized-agents.md. - Live model execution: run
a2a-bridge doctorfirst. A managed agent sandbox can lack DNS while approved host execution and computer-level auth remain healthy; repeat the exact minimal control via approved host execution before changing auth or packages. Do not trust an inherited network marker alone. Fable additionally requiresA2A_BRIDGE_ALLOW_FABLE=1; a Fable reader must mountdeploy/containers/claude-fable-settings.jsonat/root/.claude/settings.json:roalongside creds. - Concurrency: concurrent containerized runs are safe with one shared config — same repo twice or
different repos at once. Each run stamps a unique
a2a.runid into its container names (no clash) and holds an OSflocklease that marks it alive, so a peer's before-first-use recovery reaps only crashed (Dead) orphans, never a live run's containers (ADR-0025). Crash leftovers are auto-recovered before the next run and inspectable viaa2a-bridge containers list|reap. (Distinct configs are still fine, just no longer required to parallelize.)
More
docs/onboarding.md— running the bridge with your own agents, end to end.docs/containerized-agents.md— the egress-locked container setup + creds.docs/adr/— design decisions (ADR-0014 cwd, ADR-0024 warmimplementsession, ADR-0025 concurrent runs, …).
