Imported from zhiyuan-zhang0206/Ava (
AGENTS.md). Install upstream withnpx skills add zhiyuan-zhang0206/Ava. Copyright stays with the author.
Ava — minimal code-as-action agent
Small core, minimal by design. One tool (execute_code), one namespace (ava.*).
Philosophy →
Core principles
- Small core, minimal — each layer considered for removal as models improve.
- Fail fast — no fallbacks for model mistakes; use
[]not.get(), explode on unknown enums. - Don't reinvent — LangGraph, psycopg, uv; swap only when they get in the way.
- Single tool —
execute_code(code: str)+ava.*namespace = all capabilities. - Approved stable — Python 3.12, Postgres 17, Redis 8.2; upgrades require manual approval; no beta/nightly.
- English only — no raw CJK — docs, comments, prompts, error messages
in English; the only exemption is frontend i18n locale files (user ruling
2026-08-27, enforced repo-wide by
scripts/content_lint/lint_no_cjk.py).
Full elaboration: conventions/philosophy.md
Output format
One text reply + optional execute_code tool call. Lifecycle states: idle
(text only), action (text + code), restart / terminate (via ava.self).
No multi-tool dispatch, no JSON schema per capability — maximum expressiveness
with no escape hell.
OKF index →
Stack
| Layer | Choice |
|---|---|
| DB | Postgres 17 |
| Cache | Redis 8.2 |
| Framework | LangGraph (8-node self-looping graph) |
| SDK | ava (this repo) |
| Package manager | uv |
| Frontend | Next.js 16 + React 19 + Tailwind 4 + shadcn/ui |
Running
A cluster = one logical deployment; its identity IS its home path — no
cluster name (display label = home basename). A unit = one install under its
own $AVA_HOME, carrying a capability set: gateway (owns Postgres/Redis + the
HTTP gateway) and/or agent-runner (agents + the ops server) and/or
observability-station (owns the native LGTM observability backends — the
declarative form of the $AVA_HOME/lgtm-host marker); a single box carries
both gateway and agent-runner (home ~/.ava). Every cluster owns its OWN Postgres + Redis
instance under $AVA_HOME (per-cluster ports in its host-port block) plus a
PgBouncer pooler (default on — AVA_DB_URL points at it; migrations/pg_dump
dial the direct URL); isolation is home-directory isolation, so co-located
clusters share no data plane. The data plane is swappable: URLs naming a
foreign host (another machine or a SaaS provider) make the cluster treat it as
remote-managed — local instance bring-up/stop/ACL/pooler management are skipped
and degrade to reachability probes (see
docs/history/2026-08-28/connection-layer-swappable.md). Postgres
and Redis run as native processes (no Docker); gateway is POSIX-only, so a
Windows unit carries agent-runner only
(setup). Rationale + the remaining slice:
future/infra/embedded-per-cluster-data-plane.md.
Auth follows the authority boundary. AVA_CLUSTER_SECRET is the gateway's human bearer (API,
frontend login); it stays on the gateway and rotates only explicitly
(scripts/data_plane_ops/rotate_cluster_secret.py; backups use a birth-pinned passphrase). An EMPTY secret
(single-box default) leaves the API, /ops and frontend unauthenticated and binds every data-plane
listener to loopback; a set secret adds this host's reachable address for Postgres and its pooler
(Redis stays loopback, off-box inbound via the relay bridge). The internal data plane always
authenticates: Postgres and PgBouncer admit only SCRAM application logins (the OS-user administrator
and the collector's password-less monitoring role use peer on the owner-only socket), and Redis
requires its generated passwords. Application processes never hold schema-owner or admin credentials:
the owner is NOLOGIN, and each rollout's write generation — one gateway and one runner login
inheriting the NOLOGIN groups ava_gateway / ava_runner, plus one machine API token per class,
recorded in $AVA_HOME/db-authority/ — is delivered in the launch environment of the admitted runtime
(AVA_DB_URL, AVA_API_TOKEN) and, at 0600, in $AVA_HOME/run/ava-root/manifests.json until the
next start rewrites it, inert after the next fence; .env holds the credential-free endpoint. Machine
callers present their API token: the gateway admits the active generation's tokens (never a revoked
one), an ops server its generation's two. Bootstrap serves configuration only: a remote agent-runner
gets its runner login, API and telemetry tokens in a sealed bundle its start installs (ava cluster db-authority issue-unit), all shared across runner units (only the bundle's enrollment secret is per
unit), and never holds the human secret. A home born before this model (no ledger) is refused; no
conversion exists.
| Path | Role |
|---|---|
$AVA_HOME/source/ (default ~/.ava/source/) |
prod — cwd of the long-running service sessions; always the default home's cluster (its own pg 5433 / redis 6380 + prod service ports) |
~/Ava/ (this checkout) |
dev clone — worktree dev under .worktrees/<task>/ (branch from main, PR into main) (manual / agent-created) or .claude/worktrees/<task>/ (Claude Code's native worktree tool); each worktree gets its own cluster via ava start --worktree (home ~/.ava-<worktree-dir> by default), isolated db/redis/ports/sessions |
ava start is the single idempotent initialization and startup entry. Before
runtime Settings or native effects, it persists start-intent.json (home,
capabilities, checkout, ports; credentials until .env holds them); repeats retain that identity and
service selection, an interrupted one resumes, and an ambiguous home or
contradictory pointer refuses. The home is checkout-anchored: explicit AVA_HOME,
the production source path, or the checkout's .ava_home pointer (--worktree
supplies a development home); a checkout with none owns no cluster and boots
bare on a scratch home — no .env, no gateway fetch, never ~/.ava.
First start takes the machine name, capability flags and reachable host. A
remote agent-runner joins through the same entry with --gateway-url and its
capability bundle (--db-capability, the transport key in AVA_DB_CAPABILITY_KEY),
which also authenticates it; it creates no gateway or local data plane. Networked
release operations keep refusing until capability delivery is automated.
Registry records are keyed by absolute home path in ~/.ava/clusters.json
(or an explicit private AVA_CLUSTER_REGISTRY). First configuration may be
supplied with --config-file; credentials and identity survive retries.
Application processes have one supervisor: ava-root. On macOS the ancestry
is launchd -> signed permissions helper -> ava-root -> services / agent-host.
On Linux it is systemd/direct launch -> ava-root -> services / agent-host;
Linux has no helper layer. Root names the supervisor, not a privileged user.
Native data-plane custody is separate so application shutdown can retain the
database for migrations. Readiness requires a real protocol response from the
captured process generation; missing evidence cannot become success.
A checkout's own .venv/bin/ava acts on the checkout it belongs to (where its cli
source lives), not the current directory; first start runs it. The host's bare ava
(~/.local/bin/ava, linked by production start converge) is scripts/ava-launcher.sh: it
runs $AVA_HOME/ava, the CLI link every converge keeps in its own home, and refuses without
AVA_HOME — there is no default cluster. Host wiring + each plugin's scaffold() are
applied by the source-start converge phase (cli/commands/converge/host.py; standalone:
ava converge). Retained-image startup verifies its captured home and prepared artifacts;
it does not install packages, migrate or scaffold plugins.
Release preparation captures committed source, acquires hash-checked inputs and
builds a complete inactive image before any outage. ava cluster update --prepared REQUEST submits or continues the exact captured release operation through a
finite external executor. The ordinary root boot unit owns the replacement
application. Current activation supports one single-machine home (Linux, or
macOS through the home helper) with a local data plane and unchanged packaged
SQL; its stop closes terminals and schedules, exactly as PITR activation does (Linux-only). Fleet/schema
transitions and other platform adapters remain pre-cutover work; unsupported
requests refuse before draining. See
release preparation and
release transition.
uv sync # prepare the checkout dependencies and CLI; no cluster is created
ava start --worktree # initialize or resume a private development cluster
ava start # reconcile the established home and desired service roster
# --only-service NAME is an allowlist; --disable-service NAME is an exclusion
# --all-services explicitly resets selection; omitted flags retain it
ava pause # normal agent drain; preserves infrastructure, browser and persistent PTYs.
ava stop # normal agent drain, then full local stop including PTYs/browser/private pg+redis.
# --keep-infra / --keep-service retain resources; --force is explicit escalation.
# ava start resumes after readiness; agent identities and durable data survive.
ava status # check status (includes the pg/redis view)
ava cluster update --prepared /absolute/request.json
# submit or continue one captured operation; dispatch is not completion
ava cluster db-authority issue-unit --machine NAME --home UNIT_HOME --out BUNDLE
# gateway: seal one unit's capability; prints its transport key once
ava start --no-serve-gateway --serve-agent-runner --gateway-url URL --machine-name NAME --machine-host HOST --db-capability BUNDLE
# first start of a remote runner; supply AVA_DB_CAPABILITY_KEY in the
# environment; the bundle is consumed
ava cluster ls / status # list all registered clusters (label = home basename) / full multi-machine roster
ava cluster down --path PATH # stop the cluster at a home path, keep its slot (data stays on disk)
ava cluster destroy --path PATH # stop + free its slot + deregister its OS jobs (refused for ~/.ava)
Agent processes are not started directly — they always go through the
gateway via POST /api/agents (ava.agents.spawn / frontend /
scripts/start_agent.py all share this one endpoint): start gateway first,
then start agents.
Units/prod/dev paths in depth, the long-running session table, healthchecks,
and the full ops runbook: conventions/runbook.md;
dev-host inventory + secret paths: conventions/dev-setup.md.
Migrations: migrations/YYYYMMDDTHHMMSS_<kebab-name>.sql (second-precision UTC),
tracked as an applied SET keyed by name; db/schema.sql is the squashed baseline.
Every migration ships a paired .down.sql, and lossy operations go
expand-contract so any one upgrade stays reversible (scripts/content_lint/lint_migrations.py
enforces format + pairing). Adding a migration: .agents/skills/add-a-migration/SKILL.md.
Agent instruction files
This AGENTS.md is this repo's entry point for all AI coding agents; Claude Code reads it (and ui/web/AGENTS.md) directly, so there is no CLAUDE.md.
Repo skills live in .agents/skills/ (open Agent Skills standard); .ava/skills/ + .claude/skills/ link back to it, built-ins (Ava Guide, …) link in from ava_builtins/skills/.
Key docs — read on demand
Five axes, one fact per place: *.ava.okf.md next to the code = what the system is; decisions/ = why (never rewritten); future/ = plans; conventions/ = how to work;
postmortems/ = why a failure escaped — frozen incident narratives, each naming the guardrail it bought, distilled into conventions/defensive-patterns.md
(read before lifecycle / release / infra work). What the system does in time is not an axis — no run is committed; query the live one (.agents/skills/inspect-a-trace/).
Doc maintenance →
| When you need to… | Read |
|---|---|
| Understand architecture | okf/index.ava.okf.md (domain overviews + the node graph) |
| Set up dev environment | conventions/dev-setup.md |
| Run ops / deploy | conventions/runbook.md |
| Write a PR | .agents/skills/write-a-pr-description/SKILL.md |
| Understand part of the codebase interactively | ava_builtins/skills/ava-workflow/calibrate/SKILL.md |
| Follow coding conventions | conventions/python-conventions.md |
| Write SDK docstrings | conventions/sdk-docstring-discipline.md |
| Maintain docs | conventions/doc-maintenance.md |
| Know what NOT to do | conventions/non-goals.md |
| Avoid a bug class that already bit us | conventions/defensive-patterns.md (stories behind it: postmortems/) |
Change discipline
- Minimal change — but not minimal-only. Smallest diff that does the job; code that works yet hurts to change is a refactoring signal, not a reason to live with it.
- Refactoring is legitimate work — in small steps, test-backed, never bundled into an unrelated large change.
- Scope discipline — do not expand the task. An unrelated problem of the same kind found along the way is fixed in the same PR when it is a small leftover (user ruling); otherwise it must leave a trace: report the debt or hand it off — never let it evaporate.
- Unclear requirements — ask first (workflow align).
- Behavior changes are locked by a test — no tests for the sake of tests.
- Re-read the diff before committing; drop what is not necessary.
- No new dependencies or upgrades unless necessary.
Rules 1–2 and 5–7 are referenced, not restated, by
serious-engineering implementation,
serious-engineering dependency-management,
and the ava.skills.ava-code:testing discipline; rule 4's ask-first loop is workflow align.
Workflow (mandatory)
- Worktree + PR — every change in
git worktree add -b ava-<id>-<task>, merged via PR through the Trunk merge queue; direct push forbidden. Merge is not deployment; runtime rollout requires separate operator authorization and verification. Workflow → - PR description — must have file-tree diff with ★ critical paths + prose data flow. Spec →
- Tech-debt sweeps — follow
.agents/skills/ava-sweeper/(debt classes + tracker; boundary vs. lint inconventions/lint-vs-sweeper.md). - Complexity analysis — McCabe cyclomatic complexity + maintainability index via radon, ranked for refactoring. Skill →
- Local tests before push — run only targeted pytest/vitest tests and relevant eslint/tsc checks before pushing. Full test suites run in CI only; never launch a local repository-wide or full-backend test run, including for
base/changes (user ruling 2026-09-22). How to → - Git hooks — install both stages from the main clone's stable
.venv; heavy static checks run at pre-push. Install and guardrails → - CI to green, then enqueue, then clean up — poll
.venv/bin/python scripts/ci_utils.py <PR#>until all-green (fix red immediately;NO_WORKFLOW_RUNS= the suite never ran = not green), then submit with--wait --merge(submits to the Trunk merge queue; the queue verifies the combined tree that actually lands — a PR with conflicts still needs a manualgit rebase origin/mainfirst; PRs awaiting user review are never enqueued). After merge: remove the local worktree and delete the remote branch. Detail → - Commit = code + docs stable — docs go in same PR. Structure changes reconcile the co-located
*.ava.okf.md; scanconventions/+future/for stale refs.
Python conventions (quick reference)
- No
if TYPE_CHECKING:(lint-enforced). Exceptions in_TYPE_CHECKING_ALLOWED. - Structure budgets: ≤800 lines per
.py; ≤20 direct Python files/subdirectories per directory; function cc <15 (10–14 warn), nesting ≤5; packages + tests/scripts, frozen shrink-only baseline. Locality: no_-private import from outside its owning package; single-owner decisions (Postgres dial →base/db/connections.py); no path imports (sys.pathedits, file loaders) underava_builtins/, except a one-line__file__-derivedsys.pathguard that stays inside a skill's own directory — skill scripts stay thin over a package — all frozen in the same baseline. - No
print()in framework code (usebase.log.logger). - No decorative emoji in core Python.
- Import layering:
base < ava < agent < gateway < cli. Full conventions →
Communicating with the user
- No dev time estimates. Scope + trade-offs only.
- Describe current behavior; skip "used to be X, then Y" unless forwarding to
decisions/. - Clean residual old-API mentions in code/docs when found.
- Candidate next steps: list work options only — no "take a break" wrap-up suggestions. Full guide →
