Imported from Zhivex/zhivex-ai-sdk (
docs/AGENTS.md). Install upstream withnpx skills add Zhivex/zhivex-ai-sdk --skill docs. Copyright stays with the author.
Agents Guide
This guide is the adoption path for the agent-focused release of Zhivex AI SDK.
Zhivex agents are designed for server-side TypeScript applications that need portable tool-using agents across providers, resumable state, human approvals, memory, tracing, evaluations, and provider capability routing.
Related guides:
- Next.js Runner Guide: server route handlers and streaming UI shape.
- Workflows Guide: deterministic multi-step agent workflows.
- Agent Observability Guide: traces, audit records, ledgers, golden traces, and evaluations.
- Workspace Agents Guide: shell/apply-patch harnesses, approvals, and app-owned execution boundaries.
Choose The Entry Point
Use Agent for new agent code:
import { Agent } from "@zhivex-ai/sdk";
import { createOpenAI } from "@zhivex-ai/openai";
const openai = createOpenAI({
apiKey: process.env.OPENAI_API_KEY
});
const agent = new Agent({
model: openai("gpt-5"),
instructions: "Be concise and use tools when they help.",
maxSteps: 4
});
const result = await agent.run({
prompt: "Plan the next support response."
});
Use createAgent() and runAgent() when you prefer a functional API, plain object definitions, or compatibility with existing integrations:
import { createAgent, runAgent } from "@zhivex-ai/sdk";
const agent = createAgent({ model, instructions: "Keep answers short." });
const result = await runAgent(agent, { prompt: "Summarize this case." });
Use @zhivex-ai/agents instead of @zhivex-ai/sdk when you want the smaller agent-only facade. Its root contains only the stable application runtime. Import stable persistence, tracing, and evaluation helpers from @zhivex-ai/agents/ops; stable control-plane contracts from @zhivex-ai/agents/control-plane; stable live agents from @zhivex-ai/agents/realtime; and deterministic doubles from @zhivex-ai/agents/testing. The legacy @zhivex-ai/agents/beta path remains a compatibility alias and also exposes governance helpers that have not been promoted.
import { Agent } from "@zhivex-ai/agents";
import { createPostgresAgentRunStore } from "@zhivex-ai/agents/ops";
import { createAgentApprovalQueue } from "@zhivex-ai/agents/control-plane";
Use @zhivex-ai/sdk when the app also needs Runner, workflows, artifacts, embeddings, media generation, or the CLI.
Runtime Contract
Every agent run returns a serializable state:
messages: normalized messages and tool results.steps: model calls with immutable request/response snapshots and actual call timings.toolResults: executed local tool results.pendingApprovals: provider and SDK-managed local-tool approval requests that need a human decision.approvalHistory: resolved local-tool approvals, including the input digest and tool version used for replay protection.harnessandexecutionEnvironment: immutable fingerprints checked before a durable resume.compactions: replay-visible records of context summaries persisted before the next provider request.finalOutput: the validated typed result when the agent has anoutputSchema.usage: normalized token usage aggregated across every model call in the run.status:completed,waiting_approval,failed,timed_out,cancel_requested, or another production state.schemaVersionandrevision: the persistence format version and monotonic compare-and-swap revision.
Persist the state in your app, or attach an SDK run store:
import { Agent, createPostgresAgentRunStore } from "@zhivex-ai/sdk";
const agent = new Agent({
model,
store: createPostgresAgentRunStore({ client: postgresClient })
});
Built-in stores atomically claim idempotency keys before model or tool side effects. They also compare revisions before each transition, so duplicate requests share one run and stale concurrent resumes or cancellations raise ConflictError. Custom run stores must implement claimIdempotencyKey() to accept idempotent inputs. Use SQLite or Postgres for concurrent production workers; the file store is a local-development option.
For shared stores, always pass scope: { tenantId, userId?, namespace? }. It partitions runs, memory, idempotency, leases, parent indexes, and tool journals. SQL workers use renewable leases and durable checkpoints to recover expired runs. A completed tool result is reused from its journal; an indeterminate tool execution is not repeated automatically. Side-effecting tools must forward context.idempotencyKey and context.abortSignal to the external service.
Streams have bounded replay/backpressure and state has a 4 MiB default serialized limit. Request snapshots are incremental, so multi-step histories grow linearly. Telemetry and memory adapters are best effort unless hookFailurePolicy explicitly selects strict failure semantics.
Legacy states without a schema version or revision are normalized. Unknown future schema versions are rejected; use normalizeAgentRunState() or migrateAgentRunState() at application persistence boundaries.
Capsules created with createAgentCapsule() bind a canonical SHA-256 harness fingerprint to their agent definition. Resume with capsule.agent; a mismatched id, version, fingerprint, or execution-environment binding is rejected before model or tool work. The legacy migration flags in AgentRunPolicy are explicit one-time escape hatches, not compatibility defaults.
Use executionEnvironment when the application has a container, microVM, remote worker, or policy boundary that must own tool execution. The runtime acquires it once per run, authorizes every call during atomic batch preflight, reauthorizes immediately before execution, and releases it with the final status. The adapter must provide the actual isolation; the SDK contract alone is not a sandbox.
Use compaction to bound active model context without rewriting prior completed runs through AgentMemoryStore. The runtime preserves leading system messages and a recent atomic tool tail, persists the compacted messages and digests before the provider call, resets the step snapshot offset at that boundary, and exposes compactions in replay and streaming.
Tools And Safety
Tools are app-owned functions with Zod schemas:
import { Agent, applySafetyPolicyToAgent, createProductionSafetyPolicy, tool } from "@zhivex-ai/sdk";
import { z } from "zod";
const baseAgent = new Agent({
model,
maxSteps: 4,
tools: {
lookupOrder: tool({
name: "lookupOrder",
schema: z.object({ orderId: z.string() }),
execute: async ({ orderId }) => ({ orderId, status: "shipped" })
})
}
});
const agent = applySafetyPolicyToAgent(
baseAgent.toDefinition(),
createProductionSafetyPolicy()
);
Use approval policies for write, network, filesystem, code-execution, shell, payment, deployment, or other external side-effect tools. Use redaction policies before exporting traces or audit records.
Human-In-The-Loop
Provider approval waits and local tools configured with approvalMode: "interrupt" use the same resumable state pattern:
import { Agent, tool } from "@zhivex-ai/agents";
import { z } from "zod";
const agent = new Agent({
model,
tools: {
deploy: tool({
name: "deploy",
schema: z.object({ target: z.string() }),
requiresApproval: true,
approvalMode: "interrupt",
approvalVersion: "2026-07-29",
execute: async ({ target }, context) => {
return deploy(target, {
signal: context?.abortSignal,
idempotencyKey: context?.idempotencyKey
});
}
})
}
});
const waiting = await agent.run({ prompt: "Deploy staging." });
if (waiting.status === "waiting_approval") {
const approved = await agent.resume({
state: waiting.state,
approvals: waiting.state.pendingApprovals.map((request) => ({
provider: request.provider,
approvalRequestId: request.id,
approve: true
}))
});
console.log(approved.outputText);
}
A tool policy has three outcomes: { approved: true } allows execution, { approved: false } denies it, and { approved: false, approvalRequired: true } creates a resumable approval request. The runtime preflights the complete tool-call batch before executing any call, then revalidates the schema, enablement rule, guardrails, bound input digest, tool version, and optional toolApprovalSigner on resume. Local approval records stay in agent state and are not injected into provider messages.
Context is ephemeral and must be supplied again on resume. Use contextSchema to validate it and access it from policies, tool guardrails, and tool.execute through executionContext.context or context.context. Use outputSchema with outputMode: "auto" | "native" | "prompted" for a validated result.finalOutput.
For app-facing queues, import createAgentApprovalQueue() from @zhivex-ai/agents/control-plane to turn pending requests into redacted items with cryptographically random approval tokens, fingerprints, expiration, and resume URLs. Persist the queue item and token server-side, authorize the caller in the application, then call controlPlane.resumeApproval(); it validates the token, expiry, request fingerprint, and durable pending state before using store CAS to consume the approval exactly once. The SDK does not provide an HTTP authorization boundary.
Tool execution timeouts abort the AbortSignal passed as the second argument to tool.execute(input, context). Tools that perform I/O should forward context.abortSignal to their client so cancellation stops the underlying work; timeout cancellation is cooperative for tools that ignore the signal.
Child approvals are promoted to the parent as kind: "subagent" requests. The parent embeds the waiting child checkpoint in state.childRuns; after the parent decision is collected, resumeAgent() re-enters the same unresolved subagent tool and resumes the same child run. Child provider-approval messages remain isolated from the parent model transcript.
Streaming And UI
Use agent.stream() for lifecycle-aware streaming:
const stream = agent.stream({ prompt: "Draft a status update." });
for await (const text of stream.textStream) {
process.stdout.write(text);
}
const final = await stream.collect();
Use toUIAgentStreamResponse() for server routes that need UI stream chunks with agent lifecycle events, approval requests, tool progress, and final state.
Sessions
Agent runs are single-run primitives. For user-facing multi-turn apps, wrap an agent in Runner + SessionService from @zhivex-ai/sdk:
import { Agent, createFileSessionService, createRunner } from "@zhivex-ai/sdk";
const runner = createRunner({
appName: "support-copilot",
agent: new Agent({ model, instructions: "Help support agents reply." }),
sessionService: createFileSessionService({
directory: ".zhivex/sessions"
})
});
Use Postgres sessions for production serverless deployments.
Multi-Agent Patterns
Use the smallest pattern that matches the workflow:
- Handoffs: sequential ownership transfer from one agent to another.
- Subagents: model-driven delegation inside an agent loop through generated tool calls.
runAgentGroup(): deterministic fan-out from application code.- Workflows: explicit deterministic control flow with sequential, parallel, and loop steps.
Use subagents when the model should decide whether to delegate. Use workflows when your app already knows the step order.
Evaluation And Operations
Production agent work should produce inspectable artifacts:
createAgentRunSnapshot()andreplayAgentRun()for deterministic review.createAgentTraceArtifact()andsummarizeAgentTrace()for trace summaries.createAgentAuditRecord()andcreateToolAuditRecords()for redacted audit logs.createAgentRunLedger()for control-plane records that combine snapshot, replay, audit, tool audit, trace, summary, and optional cost.createAgentEvaluationFixture()andrunAgentEvaluationFixture()for regression suites.promoteAgentGoldenTrace()for turning a successful run into a regression baseline.
With the focused package, stable trace, replay, cost, provider-support, and evaluation helpers come from @zhivex-ai/agents/ops. Stable capsules, approval queues, ledgers, golden traces, capability routing, and the durable control-plane facade come from @zhivex-ai/agents/control-plane. Generic audit/governance and hosted-tool classification helpers remain Beta under @zhivex-ai/agents/beta.
The CLI in @zhivex-ai/sdk can inspect saved states and ledgers locally:
zhivex-ai agents ledger --state agent-run-state.json --out run-ledger.json
zhivex-ai agents inspect --ledger run-ledger.json
zhivex-ai agents golden --ledger run-ledger.json --name happy-path --out golden-trace.json
Release Validation
Maintainers use the agent release checklist and the canonical release workflow.
External effect reconciliation (Beta)
status: "completed" still means the agent loop finished. The additive
taskOutcome on state and output distinguishes resolved, denied, failed,
in_progress, and needs_reconciliation. The latter includes durable tool IDs,
idempotency keys and the typed INDETERMINATE_TOOL_EXECUTION diagnostic. Historical
states can omit this field. An ordinary tool error does not imply an indeterminate
external effect. resolved means the runtime has finished without unresolved
journal entries; applications must still validate their business success criteria.
A denied approval is reported as denied. Agent evaluations reject pending
reconciliation even when the technical status is completed.
Use reconcileAgentToolExecution from @zhivex-ai/sdk or
@zhivex-ai/agents/beta with FileAgentRunStore or InMemoryAgentRunStore:
const state = await reconcileAgentToolExecution({
store,
evidence, // operationId, runId, scope, toolCallId, toolName, idempotencyKey,
// exact input, confirmed output, source, proof (all JSON)
verifyEvidence: async (candidate, journal) => {
// Application-owned verification against an authenticated external ledger.
// Verify operationId, tenant, arguments, receipt integrity AND output.
return externalLedger.verifyConfirmedEffect(candidate, journal);
}
});
const result = await runAgent(agent, { state, maxSteps: state.currentStep + 2 });
const successful = result.taskOutcome?.status === "resolved" && validateBusinessResult(result);
The SDK binds evidence to the persisted run, scope, tool and input. The required
verifier authenticates the source and binds the external operation ID and result;
returning true without verification is unsafe. No evidence, rejection, conflicting
evidence or an active worker leaves the operation blocked. Evidence is stored in
the journal with its decision, timestamp and previous task outcome/output. Keep
receipts minimal and free of secrets. Identical reconciliation is idempotent.
The journal decision precedes the CAS state projection. A crash between those writes is repaired by repeating reconciliation with the same evidence. The state becomes queued; continuation is required to resolve the task. Previous steps and output remain in audit records. Exact repeats of the confirmed tool/input reuse its result within this run; use a distinct operation ID for a new intentional mutation. Other inputs are separate operations. Lease ownership is fenced at each write. SQL stores currently reject this API until they implement the same fencing capability; direct journal editing is not a supported reconciliation workflow.
Per-invocation memory opt-out
Pass memory: false to Agent.run, Agent.stream, Agent.resume, runAgent,
streamAgent, or resumeAgent to disable all AgentMemoryStore.load and save
calls. The runtime records memory: false in the run state and in declared
subagents' states, so the policy survives serialization and later resumes.
Precedence is explicit:
- A persisted
state.memory: falsealways disables memory, even when the new invocation omits the option or uses a definition with a memory adapter. - An invocation's
memory: falsedisables memory for that run and every declared descendant, overriding inherited and explicit child adapters. - When neither input nor state disables memory, existing definition defaults and subagent inheritance apply unchanged.
There is no memory: true override for an opted-out run. Start a separate fresh run
to use memory again. Definitions are not mutated, so independent invocations may
choose different policies. Resuming by serialized state, durable runId, or
idempotencyKey retains the opt-out; only the winner of a fresh idempotency claim
may read memory. A child checkpoint also retains opt-out when resumed directly
with its original definition.
Before initializing memory, the runtime claims execution with a revision check,
including when leases are unavailable or disabled. Custom stores must enforce
save(state, { expectedRevision }) atomically so a concurrent retry cannot also
initialize memory. The initialized context is checkpointed before model execution.
const result = await agent.run({ prompt: "Handle this without memory", memory: false });
// The saved state carries the policy; omission cannot re-enable memory.
const resumed = await agent.resume({ state: result.state, maxSteps: 4 });
Keep the complete SDK state when serializing checkpoints. Legacy states without
the marker preserve their historical defaults: an empty message list cannot prove
that memory was disabled. If the original policy of an unmarked checkpoint is
unknown, pass memory: false on resume to establish a disabled policy. Invalid
persisted marker values are rejected rather than interpreted as permission to
access memory.
A state may contain messages loaded before memory was disabled; opt-out prevents
new memory access but does not erase those messages. Run stores, checkpoints,
tool journals and compaction remain independent. New handoff runs and custom tools
that start agents outside the declared subagents tree are independent invocations;
pass the opt-out explicitly when they should also disable memory.
