Imported from thurlow-research/HumanOversightSystem (
docs/AGENTS.md). Install upstream withnpx skills add thurlow-research/HumanOversightSystem --skill docs. Copyright stays with the author.
[your project] — Agent Pipeline
Last updated: June 2026. Applies to the Spec 1 pilot build.
This document describes the multi-agent pipeline used to build [your project]. It covers each agent's role, model, escalation paths, and the full pipeline sequence. It is written to allow recreation in another Claude Code environment.
Design principles
The pipeline is organized around a single rule: each agent owns one concern and escalates everything else. No agent makes decisions outside its domain; disputes travel up a defined chain until they reach the right authority. This prevents agents from silently making product, architecture, or policy decisions that belong to a human or a specialized agent.
Four tiers of authority:
| Tier | Who | Decides |
|---|---|---|
| Human | You | Product vision, policy, unresolvable disputes |
| Architect | architect agent |
All technical/architectural decisions |
| PM | pm-agent |
All product/requirements decisions |
| UX Designer | ux-designer agent |
All design system decisions — tokens, component patterns, copy rules, feedback states |
Every other agent operates within the bounds set by these four. ux-designer is a peer authority to pm-agent and architect within its domain: it can extend the design pack without human involvement for additive changes, and only escalates structural brand changes upward.
Universal AI disclosure requirement
Issues
Every GitHub issue created by an AI agent must have the creating agent identified in the title:
[AI: {agent-name}] {issue-type}: {description}
Examples:
[AI: spec-red-team] spec-gap: auth flow missing rate limiting[AI: security-reviewer] security-finding: SQL injection in parking/views.py[AI: red-team/codex] red-team-finding: session fixation in invite flow[AI: claude] feat: session summary pipeline
The agent name is the agent that created the issue (e.g. spec-red-team, security-reviewer, oversight-evaluator, claude for a Claude Code session). This makes AI-created issues immediately recognisable in the issue list without having to open them.
Every AI-created issue must also include a compact footer at the end of the body:
---
*🤖 Created by `{agent-name}` | Step: {N or "session"} | Branch: `{branch}` | {YYYY-MM-DD}*
The footer is lighter than the PR disclosure block — no model ID or structured table — because issues are log entries, not merge decisions. The step/branch fields provide the build context needed for triage and to understand whether the issue is still relevant if the branch is abandoned.
This requirement applies to all projects — HOS, and any consumer repo using HOS agents. Omitting the [AI: ...] prefix or the footer is a protocol violation.
Comments (PR + issue)
Every comment an AI agent posts on a PR or issue — including review-thread replies — must open with an agent marker:
🤖 [AI: {agent-name}] …comment…
Agents run on a human's GitHub account (e.g. the HOS agent uses ScottThurlow), so without this marker an agent comment is attributed to the human and is indistinguishable from a human decision. This matters most for the needs-human ⇄ needs-ai handoff: a human's Decision: resolving a needs-ai issue must be visibly not agent text. The rule is therefore directional — the agent marks its own comments; a human comment (including a Decision:) is never prefixed by the agent. Omitting the marker is a disclosure violation, the same class as an unmarked AI-created issue or PR.
Pull Requests
Every PR opened by an AI agent — regardless of which agent, which tool, or which pipeline path — must include:
- Title prefix:
[AI: agent-name]— e.g.[AI: oversight-orchestrator]or[AI: claude] - Disclosure block as the first section of the PR body:
## 🤖 AI-Submitted Pull Request
This PR was **created and submitted by an AI agent**. A human did not manually write or submit this PR.
| | |
|---|---|
| **Agent** | `[agent-name]` |
| **Model** | `[model-id]` |
| **Submitted** | [YYYY-MM-DD] |
| **Step / context** | [build step N or session description] |
Human approval is required before merge for **MEDIUM+ risk or any protected surface** (`AGENT-IDENTITY.md §9.0`); a change within the configured `OVERSEER_CEILING` (default `HIGH`) on a non-protected surface may be approved and merged by the **overseer** — the autonomous PR-review agent (see [Autonomous Operation Layer](#autonomous-operation-layer)) — per the branch-protection rules. Either way the merge gate decides — this PR never self-merges.
- Commit trailers on MEDIUM+ changes:
AI-Model:andAI-Risk:in the commit message.
This requirement applies to every agent — oversight-orchestrator, coder, Claude Code sessions, and any other automated PR submission. Omitting it is a protocol violation. The PR template (.github/PULL_REQUEST_TEMPLATE.md) includes a reminder and the required format.
Human handoff protocol — needs-human ⇄ needs-ai
Human review is routed through labeled GitHub issues, not by anyone watching the change stream. Every agent uses this two-state handoff so neither side has to guess whose court the ball is in:
When YOU (an AI agent) need a human decision — a spec/design/policy question outside your authority, a structural or CRITICAL authorization, a loosening request, anything you cannot resolve in your domain:
- File a GitHub issue (with the AI-disclosure title prefix above) and apply the
needs-humanlabel. The label is the human's inbox. - State the decision needed precisely — ideally as a small set of options — so the human can answer in one pass. Then halt the dependent work and record the issue number where the escalating step expects it.
When picking work back up — at the start of an agentic session, and continuously in the standing daily job (#131):
- Scan open issues labeled
needs-ai— these are issues a human has decided and handed back to HOS. - Read the human's decision from the comment opener
Decision: <choice>(also honorAction:/Disposition:if present). Act on exactly that decision — do not re-litigate it. - When done, close the issue (referencing the commit/PR that implemented the decision) — EXCEPT in a repo you don't own: per
docs/CROSS-REPO-CONDUCT.md, you never close in someone else's repo; comment with completion evidence and return the label/state to the owner, who disposes. If you genuinely need more input, swap it back toneeds-humanwith a specific follow-up question — never leave it ambiguous.
The human's side of the contract (documented for completeness): to respond, remove needs-human, add needs-ai, and comment Decision: …. The human's job is to keep the needs-human queue empty; HOS's job is to keep the needs-ai queue empty.
Do not treat a bare comment as a resume signal — only the
needs-ailabel means "a human has decided, proceed." Aneeds-humanissue with new comments but no label change is still waiting on the human. This keeps the resume condition unambiguous and prevents an agent from acting on a half-finished human thought.
Pipeline overview
START
1. pm-agent — spec review, surface ambiguities, human Q&A
2. ux-designer — design pack audit against full spec; fill all gaps;
produce docs/design/UX-DESIGN-READINESS.md
3. architect — technical feasibility review, human Q&A
(reads confirmed requirements + design readiness doc)
4. ops-designer* — initial telemetry audit; produce docs/ops/TELEMETRY-SPEC.md;
submit to architect for sign-off before any build step begins;
reactive during build for ops-reviewer gaps (classifies as clarifying/
additive/structural — structural gaps require architect + human
authorization before writing)
DESIGN
5. technical-design ↔ architect — iterate until design approved
(reads confirmed requirements + ADR + design readiness doc)
SPEC REVIEW (before coding starts)
spec-red-team — adversarial spec review (using agy) to find gaps
PER FEATURE — INNER DEVELOPMENT LOOP (repeats per incremental change)
(Prompt → coder makes change → VERIFY locally → next prompt)
Rule: never issue the next prompt on a broken working tree.
Verify after every change: lint · type-check · unit tests in scope
Only when inner loop produces clean working state → move to outer pipeline
PER FEATURE — OUTER PIPELINE (once per logical change set)
6. coder [commit only after inner loop is clean]
↓ (design pack gap? → ux-designer fills it; then coder continues)
risk-assessor — score risk, validate tier, generate inspection brief
↓ (invokes prompt-fidelity, dep-mapper, risk-historian)
7. code-reviewer
↓ approved
8. security-reviewer ─┐
9. privacy-reviewer ─┤ parallel
10. ui-reviewer ─┤ → ux-designer (design pack gap or ambiguity)
11. a11y-reviewer ─┤ → ux-designer (token contrast failure / missing token)
12. ops-reviewer* ─┤ → ops-designer (telemetry spec gap) *optional
13. reliability-reviewer* ─┤ → architect (structural reliability) *optional
14. infra-reviewer ─┘ (infra files only)
↓ all approved
15. unit-test — configurable coverage + mutant score (default 80% / 75%)
↓ targets met
16. system-test — spec functional validation
↓
oversight-evaluator — check compliance & quality; recommend proceed/escalate
↓
oversight-orchestrator — open PR with AI attribution, write panel context
↓
cross-vendor panel — independent adversarial review (agy, codex, etc.)
↓
human gate — resolve panel threads; merge PR
DEPLOY
17. deploy-verify — infra checks + browser smoke tests against live prod
SUPPORT (available on demand throughout the build)
ux-designer — (1) proactive: invoked after pm-agent at project start to
audit and complete the design pack against the full spec;
(2) reactive: answers design questions and fills gaps
for coder, ui-reviewer, a11y-reviewer throughout the build
post-change-sweep — after any change: categorizes diff by domain, drives all
relevant agents in dependency order
FRAMEWORK VALIDATION (run before committing agent/doc changes)
self-reviewer — SHIPPED. adversarial review of the governance text
itself (rules, not code); one-shot seat filled by
scripts/framework/validate_self.sh
framework-validator — runs static + AI review; acts on findings
framework-setup-validator — confirms installation is correct in a new repo
doc-validator — catches the omission class of doc bug
spec-compliance-validator — pipeline vs. its own governance spec
^ these four are framework-dev only and are NOT installed into consumer
projects. self-reviewer above IS installed, because the script that
invokes it (validate_self.sh) is installed too.
The pipeline above is the work. The autonomous operation layer is the cron harness that drives it end-to-end with no human at the keyboard:
needs-ai issue
→ worker (cron) — picks the next issue, runs the pipeline above
via the specialist agents, opens nothing itself
→ specialists — pm-agent / architect / coder / reviewers / tests
→ oversight-orchestrator — opens the PR with AI attribution
→ overseer (cron) — reviews the open PR, checks tier + gates
→ AUTO_MERGE — tier ≤ OVERSEER_CEILING, non-protected, green
→ HUMAN_REQUIRED — CRITICAL / protected surface / above ceiling
→ needs-human issue, waits for human decision
See Autonomous Operation Layer for what each agent does, the interactive-vs-autonomous distinction, and escalation targets.
Autonomous Operation Layer
Two agents — worker and overseer — form HOS's autonomous operation
layer. Unlike every other agent (which runs only when an orchestrating session
calls it), these two also run unattended as cron processes launched by
bin/hos-cron, and together they build, review, and merge PRs for LOW–HIGH risk
without a human in the loop. Each has two modes; check which mode you are in
before acting, because behavior differs.
worker |
overseer |
|
|---|---|---|
| Invoked by | bin/hos-cron --role worker (autonomous); a human (interactive) |
bin/hos-cron --role overseer (autonomous); a human (interactive) |
| INTERACTIVE mode | Single human entry point for building work — routes the human's request to the right specialists | Answers questions about PR status, risk assessments, and pipeline state |
| AUTONOMOUS mode | Picks the highest-priority open needs-ai issue (priority:critical > high > medium > low; no label ⇒ low; ties broken by lowest issue number), runs the build pipeline through the specialists, and hands off to oversight-orchestrator to open a PR |
Reviews open bot PRs, applies the merge-authority matrix, and auto-merges or escalates |
| Does | Orchestrates: routing, batching, build/test/review sequencing | Evaluates: risk tier, gates, protected surfaces, branch-protection approval + merge |
| Does NOT | Never writes implementation/design/review work itself — always delegates | Never opens branches or PRs; only acts on artifacts the worker produced |
| Merge authority | None — opens PRs only | Auto-merges up to OVERSEER_CEILING (default HIGH) when checks are green; CRITICAL, protected-surface, or above-ceiling → HUMAN_REQUIRED |
| Escalation target | Files a needs-human issue (with the 4-step "How to authorize" footer) when blocked, an agent is unavailable, or a release requires authorization |
Routes to HUMAN_REQUIRED: adds the human as required reviewer (CRITICAL tier) or labels needs-human with a §8.2 escalation comment (other reasons) |
Both agents respect two safety controls every cron cycle: the activation
file (~/.hos/<repo-id>/ACTIVE) and the halt file (hos-halt). If either
check fails — or at any heartbeat (≤15m) — the agent self-terminates. Removing
or disabling the halt file is outside the overseer's authority.
Handoff between the layers and the human runs entirely through the
needs-human ⇄ needs-ai label protocol (see Human handoff
protocol): the worker labels
its PR needs-ai and hands off to the overseer; the overseer escalates by
labeling needs-human; a human responds by switching the label back to
needs-ai. For operational guidance — monitoring cron logs and intervening —
see docs/OVERSIGHT-RUNBOOK.md. For configuring this layer (the
OVERSEER_CEILING override, the active-milestone/build-plan config, and the
HOS_BOT_LOGIN identity guard), see docs/CUSTOMIZATION.md.
Agents
Model assignment is not restated per agent below — the model: frontmatter in each agent's file under .claude/agents/ is the single source of truth.
1. pm-agent — Product Manager
Invoked: At project start (first agent); anytime a product/requirements question arises during build.
Role: Owns the spec. Answers "what should the product do?" questions. Never answers implementation or architecture questions.
At project start:
Reads all five spec files (SPEC.md, SPEC-1-pilot.md, SPEC-2-subscriptions.md, SPEC-3-exchange-economy.md, DESIGN.md) and surfaces every ambiguity, gap, or underspecified behavior in a single numbered list to the human. Does not proceed until the human has answered. The confirmed answers become a requirements supplement that feeds the architect.
During build:
Answers product questions from technical-design, unit-test, system-test, and ux-designer agents, citing the spec section. If the spec is silent, escalates to the human with a precise single question.
Mid-build spec-gap response protocol:
When a spec-gap issue arrives mid-build (not just pre-coding), pm-agent must:
- Read the issue — understand what agent raised it and what decision it is blocked on
- Classify the gap (clarifying / additive / structural)
- Apply the spec update following the change-type process below
- Notify the agent that raised the issue (via the sign-off register or direct invocation) that the spec is updated and it may proceed
- Notify
architectandtechnical-designof any additive or structural change so they can assess downstream impact
Do not close a spec-gap issue without updating the spec and notifying the blocked agent.
Spec update path:
When build discoveries or human decisions require the spec to be amended, pm-agent classifies the change and applies it:
| Change type | Definition | Process |
|---|---|---|
| Clarifying | Adds precision without changing behavior — makes the implicit explicit within what the spec already requires | Update spec directly; notify architect and technical-design |
| Additive | Specifies behavior that was always implied by the spec but not yet written — filling a gap, not introducing new behavior. A new requirement that did not exist before is structural regardless of size. | Update spec; notify architect and technical-design |
| Structural | Changes existing behavior, introduces new behavior, changes scope, or introduces a new user obligation, permission, or decision point. When in doubt, treat as structural. | Draft the change, present to human for approval before writing |
Never updates the spec to rationalize code that doesn't meet the original spec — that is a spec falsification.
Escalation out: Human (spec silent or structural change required).
Escalation in: spec-gap issues from architect, technical-design, ux-designer (routed up the chain — no implementation-phase agent reaches pm-agent directly except through the design chain). Exception — spec-red-team: it operates in the spec phase, before any technical-design or code exists, so it has no implementation-design chain to route through; it creates spec-gap issues directly for pm-agent by design. To prevent an architectural/implementation-scope gap from being resolved as a pure product decision, a spec-red-team issue that pm-agent judges to be technical or architectural in scope must get architect confirmation before pm-agent resolves it (pm-agent does not unilaterally classify a technical gap as clarifying/additive).
2. architect — System Architect
Invoked: At project start (after pm-agent completes Q&A); as final escalation for technical disputes.
Role: Makes all architecture and technical decisions. Decisions are binding and final. All other agents operate within the bounds the architect sets.
At project start:
Reads the spec, the PM's confirmed requirements, and the ux-designer's docs/design/UX-DESIGN-READINESS.md. Having the complete design system upfront informs technical decisions — particularly rendering strategy, HTMX partial scope, and which views require server-side state for UI conditions. Identifies technical risks and open decisions: GiST exclusion constraint design, availability computation strategy, earned-horizon calculation placement, multi-tenant ORM scoping, PII encryption library, TOTP storage, web push architecture, Django admin extension strategy, Docker/Caddy networking. Asks the human any questions in a single list. After receiving answers, writes an Architecture Decision Record (ADR) to docs/architecture/ADR-001-pilot.md. This ADR is the input for technical-design.
Design critique loop:
Reviews every draft of the technical design document. Critiques harshly and specifically — "this is fine" is not acceptable output. Names specific failure modes and what must change. Iterates with technical-design until the design is sound.
Dispute arbitration:
When escalated disputes arrive from coder, code-reviewer, security-reviewer, or technical-design: makes a decision, states it clearly, names which agent must change course. If the dispute is actually a product question, redirects to pm-agent via spec-gap issue.
Spec-gap escalation:
When technical-design or a reviewer escalates a spec-gap that cannot be resolved at the design level (requires a product/requirements decision): create a spec-gap issue for pm-agent, halt the dependent work, and notify the escalating agent of the issue number. Do not resolve product questions within architectural authority.
Product-boundary checkpoint: Architecture decisions are final and binding — but only after the product/policy boundary is cleared. A decision that alters user-visible behavior (timing, latency, ordering, or observable failure modes — e.g. synchronous → asynchronous/queue), the cost model, deployment-topology risk, the data-retention surface, or operational obligations must route through a mandatory human/PM checkpoint before it binds. Architectural finality applies to the technical call, not to its product consequence — a decision dressed as "pure architecture" does not escape this gate. When in doubt whether a decision carries a product consequence, route it.
Startup-gap recovery: For every reactive ADR revision or post-initial-review architecture decision, the architect first asks whether it should have been settled in the initial architecture review. If so: open or annotate a startup-artifact-gap issue, update the ADR, and perform an affected-sign-offs analysis naming which prior sign-offs stand and which must re-review — design and code approved against the superseded ADR are orphaned approvals until re-checked against the revision. (Mirrors the recovery step ux-designer and ops-designer already run.)
Escalation out: Human (unresolvable after architect, or product/policy decisions, including the product-boundary checkpoint above); pm-agent via spec-gap issue (product/requirements decisions that cannot be resolved architecturally).
Escalation in: From technical-design, coder, code-reviewer, security-reviewer, privacy-reviewer, a11y-reviewer, ui-reviewer, ops-designer, ops-reviewer (spec gap unresolved after 2 cycles), reliability-reviewer (structural reliability issue), unit-test (coder refuses a testability refactor).
3. technical-design — Technical Design
Invoked: During the design phase after the ADR, and reactively whenever coder, reviewers, or test roles need the design contract clarified or find a gap in it.
Role: Translates the product spec and architectural decisions into a detailed technical specification that a coder can implement without ambiguity. Does not write application code — writes the spec for it.
Inputs (read before acting): docs/architecture/ADR-001-pilot.md, docs/pm/CONFIRMED-REQUIREMENTS.md, and docs/design/UX-DESIGN-READINESS.md. The readiness doc defines which UI states exist for each feature — technical-design uses this when specifying view contracts and HTMX partial boundaries.
Produces: docs/design/TECHNICAL-DESIGN.md, covering:
- Django model field names, types, constraints, and indexes — including GiST exclusion constraint DDL
- Multi-tenant ORM scoping strategy (custom managers, middleware)
- URL structure (
urlpatternsskeleton for every view) - View and form contracts (name, methods, auth requirement, HTMX vs. full-page — no implementation)
- Availability computation algorithm (exact query/ORM equivalent)
- Earned-horizon metric algorithm (only elapsed past hours, 180-day rolling window)
- TOTP and recovery code flow
- Notification dispatch architecture
- Admin surface design (Django admin extension vs. custom views)
- Right-to-erasure cascade
Iteration: Submits drafts to architect for critique. Does not release the design to the coder until architect approves.
During build: Answers coder's design questions. If a question reveals a gap, updates TECHNICAL-DESIGN.md and notifies the architect.
Spec-gap routing (first receiver in the chain):
When coder, security-reviewer, or privacy-reviewer escalates a spec-related gap, technical-design is the first handler:
- Gap resolvable at the implementation design level → update
TECHNICAL-DESIGN.md; notify the escalating agent - Gap requires an architectural decision → escalate to
architect - Gap requires a product/requirements decision (architect confirms) → create
spec-gapissue forpm-agent; halt the dependent work
Do not bypass this chain — agents below technical-design in the hierarchy do not create spec-gap issues directly.
Loop exit: Iteration with architect has a maximum of 5 rounds. After 5 rounds without approval, escalate to human with the iteration count, what each revision changed, and the specific point the architect has not accepted.
Startup-gap recovery: For every reactive change to the design contract, technical-design first asks whether it should have been settled in the initial technical design. If so: open or annotate a startup-artifact-gap issue, update TECHNICAL-DESIGN.md, and perform an affected-sign-offs analysis naming which prior sign-offs stand and which must re-review — code approved against the old contract is an orphaned approval until re-checked against the fix. A late design correction must not leave already-approved code unaudited against it.
Escalation out: architect (design disputes, architectural questions; max 5 rounds then human); pm-agent via spec-gap issue (product decisions architect confirms cannot be resolved at design level).
Escalation in: From coder (design questions, spec ambiguity), security-reviewer (spec doesn't cover a threat), privacy-reviewer (spec doesn't cover a compliance requirement), reliability-reviewer (undefined reliability contract), unit-test (untestable designs — behavior whose contract is ambiguous or unobservable; spec ambiguity questions go to pm-agent directly), system-test (design makes correct behavior untestable at the system level; spec interpretation questions go to pm-agent directly).
4. coder — Implementation
Invoked: After technical-design is architect-approved; iteratively per feature.
Role: Writes production Django code. Follows TECHNICAL-DESIGN.md and the ADR. Does not decide what to build.
Process:
- Reads the relevant section of
TECHNICAL-DESIGN.mdbefore writing. - Batches all questions for a section and asks
technical-designbefore writing — not mid-implementation. - Writes code following the spec's build order (§12 of SPEC-1).
- Submits to
code-reviewer. Once code-reviewer approves,security-reviewerandprivacy-reviewerrun in parallel. Does not mark a section complete until all reviewers have approved.
Key invariants enforced in code:
- Every ORM query through a tenant-scoped manager — no raw cross-tenant queries.
select_for_update()around booking creation.- Every privileged admin action writes an
AdminAuditLogentry. - No PII in logs. No secrets in source. All hex colors via CSS tokens only.
Spec-gap behavior: When implementation reveals an underdetermined spec — two valid interpretations with different behavioral consequences, or a required behavior the spec leaves implicit:
- Minor ambiguity with an obvious safer interpretation → proceed with safer choice; note explicitly in self-flag output
- Meaningful behavioral ambiguity → halt; escalate to
technical-designwith both interpretations and their implications; do not proceed on assumption
Escalation out: technical-design (implementation design questions, spec ambiguity — first receiver in the chain); ux-designer (missing design token, component class, or UX pattern); architect (disputes unresolvable at design level).
Escalation in: From code-reviewer, security-reviewer, privacy-reviewer, unit-test, system-test.
5. code-reviewer — Code Review
Invoked: After each coder pass.
Role: Reviews Django code for correctness, design adherence, and quality. Does not cover security or privacy — those are separate agents.
Checks:
- Implementation matches
TECHNICAL-DESIGN.mdexactly (names every deviation) - GiST exclusion constraint present in migration, not just asserted in model
- Availability computation and horizon metric are correct
- One-active-booking gate correctly defined
- Bookings are hour-aligned
- Every ORM query that touches tenant data goes through the scoped manager
- Django admin views are tenant-scoped
- No premature abstractions; no dead code; no hard-coded config values
- HTMX responses return partials for
HX-Request; full pages for direct navigation
Output: Every finding includes file/line, severity (blocking or suggestion), what is wrong, and what it must change to. Sends all findings in one pass. Explicit approval statement when no blocking issues.
Escalation out: technical-design (design disputes); architect (architecture disputes).
Escalation in: From coder.
6. security-reviewer — Security Review
Invoked: After code-reviewer approves (in parallel with privacy-reviewer).
Threat model: A registered resident attacking other residents or escalating privileges; an HOA admin attacking another tenant; an unauthenticated external attacker.
Checks:
- TOTP verified on every view requiring 2FA, not just at login; rate-limited
- Recovery code consumption is atomic (cannot be used twice under concurrent requests)
- Session invalidated on logout, password change, and account block; no session fixation
- Login form does not reveal whether an email exists
- Invite tokens and recovery codes use
secrets.token_urlsafe(), notrandom - Every view verifies
instance.organization == request.user.organization(IDOR prevention) - Operator console unreachable by non-superusers
- No raw SQL with string formatting; no
|safeon user-controlled data - CSRF middleware active; HTMX requests include CSRF token
- No secrets in source, templates, or logs
DEBUG = False,ALLOWED_HOSTSrestrictive, security headers set- TOTP secret stored encrypted per ADR; time window tolerance ≤ ±1 step
Output: Each finding includes severity (critical/high/medium/low), CWE class, file/function, attack scenario, and specific remediation.
Spec-gap routing: When a finding reveals the spec doesn't cover a required security property (threat model gap, missing auth requirement, unspecified data boundary):
- Do not route directly to
pm-agent— the spec gap may be resolvable at the design level - Escalate to
technical-designwith the specific gap; technical-design determines if it's an implementation design fix or requires architectural/product decisions - Continue up the chain as needed: technical-design → architect →
spec-gapissue for pm-agent
Escalation out: Architectural security flaw → architect; security policy question → pm-agent; unresolvable after those → human via Status: ESCALATED register entry. (Gaps in design/spec contract that are not policy questions go to technical-design.)
Escalation in: From coder (re-review after fixes).
7. privacy-reviewer — Privacy & GDPR
Invoked: After code-reviewer approves (in parallel with security-reviewer).
Applicable framework: GDPR (target EU hosting; possible EU data subjects in pilot). Core principle from spec: "Hash what you only verify; encrypt what you must read back; minimize collection."
PII inventory reviewed:
| Data | Required handling |
|---|---|
| Volume encryption at rest; TLS in transit | |
| Display name | Volume encryption at rest |
| Phone | Field-encrypted (reversible); optional |
| Password | Argon2 one-way hash; never recoverable |
| TOTP secret | Encrypted per ADR |
| Recovery codes | Hashed after generation; shown once only |
Checks:
- Phone field is field-encrypted, not just volume-encrypted
- No PII field is hashed instead of encrypted (breaks read-back)
- Encryption key from environment; key rotation path exists
- No PII fields beyond those the spec defines
delete_user_pii()function scrubs email/name/phone, anonymizes booking references, deletes TOTP and recovery codes, logs erasure in audit log- Consent/lawful-basis notice shown before account creation
- Any admin view rendering resident PII writes an
AdminAuditLogentry - No PII in log output;
DEBUG = Falsein production
Spec-gap routing: Same chain as security-reviewer — when a finding reveals the spec doesn't cover a required compliance property (retention policy, PII boundary, consent requirement): escalate to technical-design first; do not route directly to pm-agent.
Escalation out: Data-collection scope → pm-agent; encryption architecture → architect; retention policy → pm-agent → human; unresolvable after those → human via Status: ESCALATED register entry. (Gaps in design/spec contract that are not policy questions go to technical-design.)
Escalation in: From coder (re-review after fixes).
8. ui-reviewer — UI & Design Conformance
Invoked: After code-reviewer approves.
Role: Verifies Django templates faithfully implement the design pack (DESIGN.md + tokens.css). Not visual taste — spec compliance.
Checks:
- No hard-coded hex values; all colors via
var(--token)or provided classes --meadowand--claynot used decoratively — only for availability state signals- Spline Sans Mono (
.mono,.spot-id,.data) appears only on: spot IDs, time windows, permit-like values — not headings, body copy, or navigation - One
.btn-primaryper view maximum .badge-availableand.badge-bookedinclude text labels, not color only.baymotif used only for: available spot framing, empty states, or logo — not as generic borders- Voice/tone: plain active labels ("Book this spot", not "Submit booking request"); sentence case; no "monetize", "asset", "module", "leverage"
- Error messages explain what to do next ("No spots open then. Try a wider window.")
- Empty states invite action ("List the first spot in your building.")
Escalation out: ux-designer (design pack gap — missing token, component, copy pattern); coder (implementation bugs); architect (shared architectural dependency); human (design-intent ambiguity or unresolved loops after 2 cycles).
Escalation in: From coder (re-review after fixes); from ux-designer (re-review notification after design pack extension).
Loop protocol: When escalating a gap to ux-designer, state the specific missing element. After ux-designer fills the gap and notifies, re-review against the updated design pack. Maximum 2 cycles; escalate to human if unresolved.
9. ux-designer — UX Design Authority
Invoked: At project start (after pm-agent completes Q&A); reactively throughout the build whenever any agent encounters a design pack gap.
Role: Owns and extends the design pack (DESIGN.md, tokens.css, style-guide.html, feedback-states.html). Answers design questions directly rather than escalating to the human. The design pack is a living specification — this agent completes it at the outset and fills gaps as new features are built.
At project start:
Reads the full spec (SPEC-1-pilot.md) and the pm-agent's confirmed Q&A output. Walks every user-visible feature in the spec and checks whether the design pack covers all required UI states: spot card states, booking gate-blocked states, authentication screens, onboarding flows, notification copy, leaderboard/gamification display, HOA and operator portal views, error and empty states, right-to-erasure.
For each gap found: fills it directly (additive/clarifying) or surfaces to the human (structural). After all gaps are filled, writes docs/design/UX-DESIGN-READINESS.md — a feature-by-feature coverage table, a log of every addition made, any open structural questions and their answers. The architect and technical-design agent read this document before starting their own work.
During build (reactive):
| Invoker | Reason |
|---|---|
coder |
Missing token or component class during template implementation |
ui-reviewer |
Gap found during template review (missing class, undocumented pattern) |
a11y-reviewer |
Token fails contrast check; accessible alternative needed |
technical-design |
New feature needs a UX pattern spec before technical design is written |
pm-agent |
Product decision has UX implications |
Change classification (mirrors pm-agent's taxonomy):
| Type | Definition | Process |
|---|---|---|
| Clarifying | Adds precision to an existing rule without changing meaning | Updates design pack directly |
| Additive | New token, component variant, or copy pattern | Adds to design pack; consults pm-agent if it affects a user flow; notifies a11y-reviewer for new color tokens |
| Structural | Changes a core color, removes a component, or changes the design brief. Also structural (per ux-designer.md): any change that introduces a new user decision point, new blocked/permission state, new completion criterion, or new step in a user flow — even if it feels small. When in doubt, treat as structural. |
Presents to human for approval before writing |
Additive is the normal operating mode. Missing error color palette, a new badge variant, a copy pattern for an empty state — all handled without human involvement. But a change that alters a user flow — a new confirmation step, a new blocked state, a new completion criterion — is structural, not additive, regardless of how small the visual change is; it requires human approval. This doc must not use a narrower structural definition than .claude/agents/ux-designer.md (the authoritative source); the oversight-evaluator independently re-derives the mechanical structural signatures (contract §2a) so a flow change mislabeled additive is caught.
After extending the design pack: Notifies the invoking agent with the exact change; notifies a11y-reviewer for new color tokens; notifies ui-reviewer so it can re-check template conformance. Appends a one-line entry to the ## Change log section of DESIGN.md.
Escalation out: brand-direction changes, structural paradigm changes, or structural UX changes (new user decision points, blocked/permission states, completion criteria, or flow steps) → human; out-of-scope addition or flow-behavior question → pm-agent (then human if out of scope); shared architectural dependency → architect; unresolved → human.
Escalation in: From pm-agent (at project start); from coder, ui-reviewer, a11y-reviewer, technical-design, pm-agent (during build).
Loop exit: ui-reviewer and a11y-reviewer escalation cycles have a 2-round maximum. After 2 cycles without resolution, escalate to human.
10. a11y-reviewer — Accessibility
Invoked: After code-reviewer approves.
Compliance target: WCAG 2.1 AA. Treats the design pack's quality floor as a build gate: keyboard focus, color never the only signal, prefers-reduced-motion, mobile responsiveness, WCAG AA contrast.
Audit approach: Lighthouse audit via Chrome DevTools MCP on each primary view (if dev server is running); plus static template analysis (grep for missing alt, unlabeled inputs, tabindex="-1" on interactive elements) in all cases.
Key checks:
- Every interactive element reachable by Tab in logical order
- Focus ring visible on every focused element; not overridden anywhere
.badge-available/.badge-bookedhave text labels, not color only--meadow-ink(not--meadow) used for colored text on light backgrounds; same for clay--slateon--canvasmeets 4.5:1 contrast ratio- No animations outside
@media (prefers-reduced-motion: reduce)guard - Every
<input>has a programmatic<label>(not just placeholder) - Error messages associated via
aria-describedby - Touch targets ≥ 44×44px; no horizontal scroll at 375px viewport
Escalation out: accessible-token/pattern gap → ux-designer (2-cycle cap → human); design-system ambiguity ux-designer cannot settle → human; implementation bug or non-token CSS fix → coder; unresolved → human.
Escalation in: From coder (re-review after fixes).
11. ops-designer — Observability Authority (optional — projects with ops complexity)
Invoked: At project start, after architect completes the ADR. Submits the completed TELEMETRY-SPEC.md to architect for sign-off before any build step begins. Reactive during the build when ops-reviewer escalates a spec gap.
Role: Authors and maintains docs/ops/TELEMETRY-SPEC.md — the observability contract that ops-reviewer enforces. Covers structured logging conventions, metric naming, distributed tracing requirements, health check requirements per dependency type, and dashboard/alerting intent. Does not implement instrumentation — records the contract for the build to follow. During the build, classifies reactive spec-gap requests as clarifying (interpretation only), additive (new spec entry, no architecture change), or structural (new external dependency, trust-boundary change, or architecture change). Structural changes require architect sign-off and human authorization before the spec is updated.
Escalation out: architect (new external dependency, trust boundary, or observability-architecture change; 2-round consultation cap then human); pm-agent (product-scope question surfaced while gap-filling); human (structural change unresolvable after 2 architect rounds, or unresolvable escalation — requires human authorization artifact before spec update).
Escalation in: From ops-reviewer (spec gaps); architect (observability ADR inputs).
N/A for: CLI tools, libraries, scripts, or projects without background jobs, external integrations, or multi-service architecture.
12. ops-reviewer — Observability Review (optional — projects with ops complexity)
Invoked: After code-reviewer approves, in parallel with security-reviewer and privacy-reviewer, when changes introduce new operations, external calls, background jobs, async tasks, or failure paths.
Role: Reviews code changes for conformance with docs/ops/TELEMETRY-SPEC.md. Asks: "Can you tell what's happening and debug it?" Withholds sign-off on silent failures and spec violations. Escalates spec gaps to ops-designer (not coder — coder cannot be held to an unspecified requirement). Loop exit: after 2 failed re-reviews for the same gap, escalate to architect.
Scope boundary: Does NOT cover security audit logging (security-reviewer), GDPR/data retention logging (privacy-reviewer), or deployment config (infra-reviewer).
Escalation out: Spec gap → ops-designer; after more than 2 unresolved ops-designer cycles → architect → human; any unresolvable issue → human via Status: ESCALATED register entry.
Escalation in: From post-change-sweep, direct invocation.
N/A for: Same projects as ops-designer. If docs/ops/TELEMETRY-SPEC.md is absent on a project with ops complexity, block and invoke ops-designer first.
13. reliability-reviewer — Resilience Review (optional — projects with external connections)
Invoked: After code-reviewer approves, in parallel with security-reviewer and ops-reviewer, when changes introduce or modify outbound connections (DB queries, HTTP calls, queue operations, cache reads/writes).
Role: Reviews code for resilience against external dependency failures. Asks: "What happens when an outbound connection fails, times out, or returns an error?" Distinct from ops-reviewer (observability) and security-reviewer (security). Complementary: a system can be well-observed and secure but still brittle.
Review dimensions: timeouts on all outbound connections, retry with exponential backoff and limit, no tight retry loops, non-idempotent operations protected from accidental retry, graceful degradation / fallback, no unbounded waits (thread pools, connection pools, queue consumers).
Escalation out: technical-design (retry/timeout policy not specified in technical-design — first receiver; does NOT create a spec-gap issue directly, same as security-reviewer/privacy-reviewer); architect (structural reliability design — sync vs async, circuit-breaker architecture); ops-reviewer (telemetry gaps on failure paths — reliability-reviewer notes these for ops-reviewer and does NOT block reliability sign-off on that lane). technical-design revises the contract or routes product-policy questions onward.
Escalation in: From coder (re-review after fixes).
N/A for: CLI tools, libraries, or any project without outbound connections to external dependencies.
14. infra-reviewer — Infrastructure Review
Invoked: Independently of code-reviewer (it reviews infra config, not application code) when infrastructure files are modified: Compose, Caddyfile, backup scripts, .env.example. An infra-only diff runs infra-reviewer directly; code-reviewer returns N/A.
Role: Reviews deployment configuration against the spec's §2 deployment requirements. Does not review application code.
Checks:
- All three services present (
web,db,caddy); all withrestart: unless-stopped - DB port not published to host; DB on internal network only
- Postgres data on a named volume, not a host-mount path
- No secrets in
environment:blocks; all via.env/${VAR}references - Caddy: canonical domain via DNS-01; HOA alias via HTTP-01; no
tls internal - Both canonical and HOA alias in
ALLOWED_HOSTS .env.examplecontains all required variables;DEBUGdefaults toFalse;DATABASE_URLuses internal service namepg_dumpbackup script exists; output to NAS/external volume; retention policy present- Portability: can the stack move to a new host by copying
.env+ restoringpg_dump+ repointing CNAME?
Escalation out: Suspicious application-config value → coder / technical-design; architecture toolchain choice → architect; deployment policy → human; unresolved → human via Status: ESCALATED register entry.
Escalation in: From coder, deploy-verify (infra failures post-deploy).
15. unit-test — Unit Tests
Invoked: After all reviewers (code-reviewer, security-reviewer, privacy-reviewer, ui-reviewer, a11y-reviewer, infra-reviewer) have approved.
Gates (both must be met before advancing):
- Code coverage ≥ 80% (
coverage run+coverage report) - Mutant score ≥ 75% killed (
mutmut run— Python mutation testing)
Priority test areas:
- Booking gate logic — all three gates tested at boundaries (horizon, one-active-booking, DB overlap constraint triggered directly)
- Earned-horizon metric — elapsed hours only, 180-day window, formula, cold-start grace, zero-history baseline
- Availability computation — window splitting, clipping, fully-booked window
- Model constraints — hour-aligned bookings, duration cap, organization scoping
- Auth flows — TOTP valid/invalid/expired/reused; recovery code single-use; invite token single-use/expiry
- Right-to-erasure — all PII scrubbed, bookings anonymized, codes deleted
- Admin audit log — every privileged action writes exactly one entry with all required fields
Tooling: pytest-django, coverage, mutmut, factory_boy, freezegun (for time-dependent tests).
Escalation out: technical-design (untestable designs — behavior whose contract is ambiguous or unobservable); pm-agent (spec ambiguities — what the product should do); architect (coder refuses testability refactor).
Escalation in: From coder (fixes that re-run tests).
16. system-test — System & Functional Tests
Invoked: After unit-test meets both targets.
Role: Validates the application meets the spec's functional requirements. Tests are based on the spec, not the code. Uses Django test client (not Selenium) against a real test database.
Covers every primary flow from SPEC-1 §11:
- Full booking flow: search → horizon gate → one-active-booking gate → overlap gate → confirm → notifications
- Listing flow: availability window creation, elapsed hours accumulation (with
freezegun) - Cancellation/release: borrower pre-start, early release, owner-cancel with penalty
- Onboarding Mode A (invite): single-use link, TOTP enrollment, recovery codes
- Onboarding Mode B (approve): pending → approved → active
- Authentication: TOTP required; recovery code consumption; locked-out sessions
- Earned-horizon advancement: baseline, cold-start grace, formula verification
- HOA portal tenant isolation: cannot see another building's residents
- Operator console: full cross-tenant access; HOA admin cannot reach it
- Right-to-erasure: PII scrubbed, bookings anonymized, audit log entry
- Admin audit log: admin-cancel, PII access, block/unblock all logged
When a test fails:
- Code bug (code doesn't match design) → report to
coderwith test name, expected vs. actual, spec citation - Spec gap / interpretation dispute (spec does not define the behavior clearly) → escalate to
pm-agentwith the exact behavior in question, the two possible interpretations, which one the test assumes, and the spec section reference;pm-agentescalates to the human if the spec is genuinely silent - Design makes correct behavior untestable at the system level →
technical-design, which makes the behavior explicit and testable
Escalation out: pm-agent (spec interpretation, ambiguity — what the product should do); technical-design (design makes behavior untestable); coder (code bugs); persistent failure past 5-round cap → file bug issue, write Status: ESCALATED, escalate to architect (unresolved → human).
Escalation in: From coder (fixes).
17. deploy-verify — Deployment Verification & Production Smoke Tests
Note:
deploy-verifyis a consumer-project agent — it is not shipped in the HOS framework source. Consumer projects generate this agent during./bootstrap/hos_install.shbased on their deployment stack. It is listed inscripts/framework/config.shasEXTERNAL_AGENTSso HOS's own static checker correctly skips the agent-file existence check.
Invoked: After docker compose up on opus.[your-domain].
Role: Verifies the production instance is correctly configured and functionally operational. Last gate before announcing a deployment successful.
Phase 1 — Infrastructure:
Remote checks (SSH to parkshare-agent@opus.[your-domain]): Docker services up and healthy, backup file exists and is recent (< 48h old).
Local checks (run from wherever Claude Code is): DNS resolution for canonical URL and HOA alias, TLS certificate valid and not expiring within 30 days, HTTP security headers present (Strict-Transport-Security, X-Frame-Options, X-Content-Type-Options, Content-Security-Policy), DB port 5432 not reachable externally, HTTP → HTTPS redirect working.
Requires three environment variables in .env: AGENT_SSH_KEY (path to parkshare-agent private key), AGENT_COMPOSE_PATH (path to compose file on opus), AGENT_BACKUP_DIR (path to backup directory on opus).
Phase 2 — Browser smoke tests (Chrome DevTools MCP):
- App loads; login form present; no console errors
- Hanken Grotesk font loaded;
tokens.cssloaded (--pineCSS variable defined) - Invalid login returns error state, not 500 or Django debug page
- HOA alias redirects to HTTPS without certificate error
- PWA manifest served as valid JSON with required fields
tokens.cssstatic file returns 200- Django admin login page loads
Phase 3 — Backup verification:
- Backup cron is registered
- At least one backup file exists and is non-zero
Output: Structured pass/fail table per check, overall PASS/FAIL, and specific remediation steps for any failures.
Escalation out: infra-reviewer + human immediately (infrastructure failures); coder + system-test (functional failures); human immediately (missing backups — deployment is not complete without verified backup).
18. spec-red-team — Spec Red-Team
Invoked: Before coding begins on a build step (after the technical design is approved).
Role: Adversarially reviews spec sections before coding. Finds gaming vectors, contradictions, implicit assumptions, and missing edge cases.
Process:
- Formulates 5–10 adversarial questions based on the spec section and technical design.
- Invokes
agy(Gemini) with an adversarial prompt to ensure vendor-independent analysis. - Reviews findings and creates
spec-gapGitHub issues for genuine problems.
Escalation out: pm-agent (to resolve spec gaps).
Escalation in: None.
19. risk-assessor — Risk Assessor
Invoked: After the coder completes a build step, before the internal review chain starts.
Role: Evaluates code changes to establish a validated risk tier and produce a ranked inspection brief for reviewers.
Constraints:
- Can only raise the coder's self-declared risk tier, never lower it (unless a human tier override exists).
- Must produce a ranked inspection brief.
Process:
- Applies deterministic floor rules (e.g. auth/PII changes force HIGH tier, booking gate forces CRITICAL).
- Runs static and IP validators (
run_validators.sh,prompt_audit_risk.py,ip_check.py). - For MEDIUM+ steps, invokes the
prompt-fidelitysubagent (semantic prompt-vs-code comparison). For HIGH+ steps, invokes thedep-mapperandrisk-historiansubagents. At CRITICAL, additionally reads the relevant spec section and performs direct spec-code fidelity checks (does the code implement what the spec says?). Note:prompt-fidelityis NYI — it returns a non-blocking coverage gap until the semantic comparison is implemented. - Synthesizes risk scores to determine the final validated tier.
- Produces a ranked inspection brief. Writes risk-assessment.md and required-reviewers.md (oversight-evaluator unions this list with the manifest so dynamic risk findings can add but never remove required reviewers).
Escalation out: Writes risk-assessment.md and required-reviewers.md; records unresolved blocking findings for evaluator. For HIGH+ dep-mapper LOW confidence without suspension, escalates to human via a blocking finding requiring either stack-specific dep-mapper override or human-authorized suspension. Escalation in: None.
20. risk-historian — Historical Risk Analyst
Invoked: Subagent of risk-assessor (runs only at HIGH+).
Role: Queries GitHub issues and git logs to build a historical risk profile of changed files.
Process:
- Queries GitHub issues matching specific risk labels (e.g., bug, security-finding, design-concern, spec-gap), excluding
duplicate-labeled issues (a re-filed finding must not inflate density). - Analyzes git log churn (commits in the last 90 days) and fix commit density (commits matching fix/bug/error in the last 180 days).
- Returns raw counts and issue references plus a
Data confidencefield — it does not classify risk.risk-assessorperforms all risk classification from this raw data (per DECISIONS.md D31; a Haiku/Sonnet retriever must not make the judgment call). A future implementation must not add LOW/MEDIUM/HIGH classification here — that would violate the subagent boundary.
Escalation out: risk-assessor (reports raw findings).
Escalation in: risk-assessor.
21. dep-mapper — Dependency Mapper
Invoked: Subagent of risk-assessor (runs only at HIGH+).
Role: Maps the project's dependency graph for changed files (imports, references, and framework wiring) to assess the blast radius.
Process:
- Checks direct imports and references across the codebase.
- Identifies framework-level implicit wiring (signals, events, middleware, views, templates).
- Classifies the blast radius category and applies risk multipliers.
- The generic mapper must report Data confidence HIGH|LOW. LOW is required when framework wiring or outward references cannot be traced; at HIGH+ risk-assessor records this as a blocking finding unless dep-mapper is human-suspended.
Escalation out: risk-assessor (reports findings and data confidence).
Escalation in: risk-assessor.
22. prompt-fidelity — Prompt Fidelity Validator
Invoked: Subagent of risk-assessor (runs only at MEDIUM+).
Role: Performs semantic comparison of prompt artifacts against generated code to verify faithful implementation.
Status: Designed and stubbed — NYI (Not Yet Implemented). The stub returns Status: NYI and does not block or escalate; risk-assessor records the coverage gap in the inspection brief. Note that NYI because semantic comparison is not implemented is a non-blocking coverage gap, but a missing prompt artifact on a MEDIUM+ step is a compliance issue surfaced under Human Review Required and enforced by evaluator prompt-artifact compliance checks. It does not perform a best-effort comparison — a stub that returns plausible results would create false confidence (see the agent file's stub-behavior rule).
Process:
- Verifies positive fidelity (implements all requirements).
- Verifies negative fidelity (adheres to negative constraints).
- Catches scope creep and prompt-code discrepancies.
Escalation out: Fidelity gaps → risk-assessor; missing prompt artifact → report inability to assess and surface to human/risk-assessor as a MEDIUM+ compliance gap. Distinguish NYI semantic comparison from missing artifact.
Escalation in: risk-assessor.
23. oversight-evaluator — Oversight Evaluator
Invoked: After all internal reviewers approve a build step and system tests pass.
Role: Evaluates compliance and quality of the build step review process.
Process:
- Phase 1 (Compliance): Verifies risk-assessment presence/scope and unresolved blocking findings, unions manifest and risk-assessor required reviewers, validates register entries including N/A reasons, re-derives N/A and structural-override waivers from the diff, enforces second-review behavior for MEDIUM+, prompt-artifact compliance, gate suspension limits, and the effective human gate.
- Phase 2 (Quality): Reviews convergence failures (long reviewer loops, overrides), resolved critical findings, confidence gaps, second review findings.
- Produces a final recommendation (
PROCEED,CONDITIONAL_PROCEED, orESCALATE).
Escalation out: oversight-orchestrator (via recommendation output).
Escalation in: None.
24. oversight-orchestrator — Oversight Orchestrator
Invoked: After oversight-evaluator produces its recommendation.
Role: Acts on the evaluator's recommendation to open PRs, prepare p
*Truncated - read the full file at https://github.com/thurlow-research/HumanOversightSystem/blob/47e39b9e4480b78b0e18d202392ad51c003e726d/docs/AGENTS.m
