Imported from zaxbysauce/zmem (
skills/memory/SKILL.md). Install upstream withnpx skills add zaxbysauce/zmem --skill memory. Copyright stays with the author.
Memory (ZMem)
ZCode's memory system has three tiers:
- Tier 0 — Core:
core.md(user-level, in the canonical shared store dir)<repo>/AGENTS.md(project-level on ZCode). Auto-injected every session by the SessionStart hook. Editcore.mddirectly for stable rules/preferences. Keep <2KB.
- Tier 2 — Semantic:
store.sqlite(in the plugin data dir). Cross-task lessons, facts, conventions, preferences. Operated by this skill viascripts/store.py. - Tier 4 — Procedural: the skills library. (Extended later with evals + index.)
Finding store.py
The SessionStart hook injects the absolute path to store.py into context each
session (look for # Memory skill: invoke "...store.py" <subcommand>). Use that
exact path. On Windows it will be a Windows-format path like
C:\Users\...\plugins\data\zmem@...\skills\memory\scripts\store.py
(the ... segments are elision placeholders, not real paths).
If you cannot find the injected path, the script is at the plugin root under
skills/memory/scripts/store.py.
When to use
- Before a non-trivial task:
recallrelevant past lessons. High-precision-first: if a retrieved lesson does not clearly apply, ignore it. - After a failure with a generalizable lesson:
adda lesson grounded in an external signal (test/compile/lint/reviewer/user). The reflection Stop hook will prompt you automatically when failures are detected. - When you learn a stable fact or convention:
addit. - When a memory is stale/wrong:
supersedeit for a general tombstone; use the dedicatedinvalidatecommand when a fact is no longer true (it REQUIRES a reason so the correction is auditable); useupdateto revise a memory append-only (the old row is tombstoned, a new live row links back viaupdate_of, and point-in-time--as-ofrecall still sees the old content).
Commands
Store commands run python <store.py path> <subcommand>; doctor is the one
separate diagnostic script shown below. On Windows use python
(NOT python3 — that is a Windows Store stub).
doctor — read-only install diagnostics
python <doctor.py path> [--project <repo>] [--repo-root <zmem-repo>] [--format human|json|both]
Read-only preflight for cutover and operator debugging. It never mutates the store or host config. Checks:
- resolved store path and split-brain env/config risk
- local/non-OneDrive store path safety
- Python version (supported floor 3.11) + SQLite FTS5
- Node and a usable Git Bash/Cygwin shell on Windows
- best-effort read/write access to the store path
- schema compatibility against current v14
- v9 append-only lineage columns present (
valid_until/update_of/taint) - v10 entity identity tables present and non-vacuous (
entity/entity_alias/memory_entity); inspect deeper withstore.py entity-list - v11 link surface present (
memory_linktable +memory.trust_scorein range [0,1]); inspect deeper withstore.py links --id <uuid> - v13 episode storage present with counts (
episode-tablescheck), and the MCP token scope advisory (mcp-tokencheck: warnsunscoped_token: trueon full-access operator tokens, never reports the token value) - Claude/Codex native-memory conflicts via read-only config inspection
- ZCode native memory (
zcode-native-memorycheck, issue #185):~/.zcode/v2/setting.jsonmemoryEnabled: truefails cutover, explicitfalsepasses; unreadable/missing setting warns — never auto-edited - host install-skew (issue #185):
duplicate-install(fail) when more than one enabled user-scope zmem install is registered for a host;marketplace-skew(warn) when an installed cache version differs from the marketplace version its registry entry points at;project-pin(warn) when a project-scoped zmem pin is behind the enabled user-scope install. Registries are read through the #184 strict codecs; missing registries skip and malformed ones warn — doctor never edits host state - Codex manifest hook trust (
untrusted-hookcheck, issue #185): compares the pre-approval events the repo manifest registers (SessionStart, PreToolUse) with the events the Codex config records trusted for the repo (falling back to the box-wide union when no entry names this repo — on a multi-repo box another repo's approval can stand in, so treat a pass as inventory, not proof). Missing registered events warnuntrusted-hook <ids>; reapproval is always manual - orphan-store inventory (issue #185): warns with
schema=/rows=for every non-canonical SQLite store on the known host paths (plugin-data env dirs,~/.zcode/memory/store.sqlite); inspect then merge withpromote-store --from <path>— doctor never deletes or migrates - canonical namespace for the provided project
- host surface presence (Claude plugin, ZCode plugin, memory skill; repo-local Codex adapter files are optional until that lane exists)
- served-tree drift vs
release-manifest.json(served-driftcheck, issue #107):matched/drifted(warn — first 10 differing paths + both digests)/unknown-skip when the tree predates 0.17.0 and carries no manifest, andunknown-WARN when a manifest is present but corrupt (failed its algorithm/digest integrity gate — the served mirror itself is damaged). Log-only; a drifted session start also appends onezmem-driftline tozmem-bg.logand shows the operator a systemMessage — never model context.
Use it before first install, before cutover, and after any store-path or hook surface change.
doctor --miss-rate — the miss-rate join + false-injection counter (issues #94, #129)
python <doctor.py> --miss-rate --store <snapshot-store.sqlite> [--miss-db PATH]
[--miss-transcripts GLOB ...] [--miss-bg-log PATH]
[--miss-window-before 1800] [--miss-window-after 300]
[--miss-limit 200] [--miss-min-overlap 2] [--miss-verbose]
[--format json]
Opt-in check that measures the miss rate — "a failure occurred in a session,
a matching memory existed in the store, and nothing surfaced at that moment".
It joins mined tool failures (ZCode episodic db via --miss-db, transcript
JSONL globs via --miss-transcripts) against a snapshot of the store
(recall with telemetry fully disabled, link_hops=0) and the bg log's
injection decision lines — BOTH shapes: writer A's reason=injected and
the session-start writer's status=injected lines (that writer has never
carried reason=). Definitions (pinned):
missed— failure + floor-passing store match + no injection of a matched row in the window ⇒ counts toward the miss rate.capture-gap— failure + NO store match (write-side problem, counted separately).surfaced (sid)— an injected line whosesid=proves it is this session's decision carries a matched row id.surfaced (legacy)— same evidence on a pre-#94 sid-less line or asid=unknownline (window+overlap attribution only;miss_rate_strict_sidexcludes it).no-query— nothing derivable for the failure (no recovered operation and no ops ring): unmeasurable, excluded from every rate. On boxes where the ops-ring lane is not yet deployed this bucket can dominate until rings accumulate.
Issue #129 adds the OTHER gate direction to the same report: the
false-injection rate — of the injected decision lines, how many were
NEVER referenced by any later same-session operation, prompt, or captured
failure. A line counts as USED when a later reference (mined failure
operation/error text, the session's ops-ring events, or transcript
prompts) contains one of its memory ids literally, or shares >=
--miss-min-overlap (default 2) distinct ops tokens with the injected
row's content. Rates print overall AND per moment (session_start,
user_prompt, pretool, subagent, precompact; pre-#129 lines
without moment= report under legacy with a caveat). Each decision
LINE counts once — the same row re-injected at two moments is two
denominator lines, never double-counted.
Decision telemetry lives in <data dir>/zmem-decisions.log (issue #129
split): maintenance output and zmem-drift lines stay in zmem-bg.log.
Both logs rotate instead of truncating — ZMEM_BG_LOG_MAX_BYTES (default
262144) caps the active file and ZMEM_LOG_ROTATIONS (default 3) keeps
bounded segments named <name>.1..<name>.N, each stamped with a
# zmem-seq= marker; the join and the counter read rotated segments.
Legacy deployments with only zmem-bg.log keep working: readers fall
back to it when no decisions log exists, and an explicit --miss-bg-log
wins over both.
REFUSES to run without an explicit --store, and refuses the
host-default store even when given explicitly — the join reads session
data, so snapshot first: copy store.sqlite AND any
store.sqlite-wal/-shm beside it into a temp dir (plus
zmem-decisions.log and its rotated .1...N segments, zmem-bg.log,
and the ops/ ring dir when present), then pass that path. exit 1 from
--miss-rate most often means exactly that:
snapshot the store and re-run with --store; exit 1 from
other checks means a real diagnostic failed.
Read-only everywhere: the
store/db open mode=ro, the report never writes, and a missing/old store is
an error (never created, never migrated). Memory content stays out of the
default output (ids/namespaces only; --miss-verbose adds a short preview).
recall — surface relevant memories (high-precision)
python <store.py> recall --query "<query>" [--namespace NS] [--limit 5]
[--link-hops 0|1] [--link-budget N]
[--include-global] [--global-limit 3] [--hybrid]
[--no-hybrid] [--no-mmr] [--no-bump]
[--as-of ISO-8601] [--no-unfold] [--json]
python <store.py> recall --query "<query>" --explain [--target ID|FRAGMENT]
[--json]
Returns live (non-superseded) memories matching the query, filtered by confidence
floor (>=0.25) and namespace — unless --as-of ISO-8601 is given, in which case
it is a point-in-time read that returns rows valid at that instant (a row
superseded after that instant may be included; a row created after it is not;
see the --as-of notes below). Prefer --namespace project:<basename> to scope to
the current project; use user:global for cross-project.
Retrieval debugger: --explain / --target (issue #82)
recall --explain re-runs the REAL pipeline (same lanes, same floors, same
merge) and prints one blameline per verdict explaining why a row did or did
not surface. It is an operator diagnostic in the same spirit as doctor.py:
ZERO writes (the telemetry path is never reached — stricter than
no_telemetry), it never unfolds, it never fails a recall (a thrown tracer
degrades to the explain_unavailable verdict and the results still print).
--target takes a memory id (full or unambiguous prefix) or a content
fragment (case-insensitive substring, then token overlap >= 0.7); multiple
matches produce one verdict per id, never a guess. With --json the read
envelope gains an explain object: query, query_shape (the
normalized terms + exact FTS MATCH expression, #112), target, the
effective settings (namespace, limit, include_global,
global_limit, no_mmr, no_bump, as_of, hybrid), and
verdicts[] with id/reason/rank/score/detail.
Verdict reasons are a CLOSED set (EXPLAIN_REASONS in storelib/recall.py):
| reason | meaning |
|---|---|
found |
in the presented result set at rank (1-based) |
below_limit |
retrieved and scored but beyond --limit |
below_floor |
the row's confidence is below the effective floor (the recall min_confidence parameter, default CONFIDENCE_FLOOR) |
omitted_injection |
would have ranked; dropped because no_bump and prompt-injection risk |
omitted_untrusted_web |
would have ranked; dropped because no_bump and taint=untrusted_web |
namespace |
lives outside the query's expanded namespace set (and --include-global did not admit it) |
superseded |
live-only recall and superseded_at is set (detail.successor_id names the live replacement when one exists) |
not_valid_at_as_of |
--as-of set and the validity interval does not contain it |
vec_lane_miss |
--as-of + hybrid: the row was valid then, but the vec KNN pool is live-rows-only (the SKILL.md as-of caveat — this is why history-on-hooks is not "just pass --as-of") |
not_in_pool |
not in the FTS, vec, or entity candidate pool at over-fetch depth |
not_in_db |
no row matched --target (detail.neighbors lists up to 5 nearest live rows by token overlap) |
explain_unavailable |
the tracer threw; results still returned (fail-open) |
For an injection-path explanation, combine --for-injection --explain. This
is a read-only replay: it follows the same omit, selective-inject,
score-margin, and token-budget stages as a passive injection, but never writes
surface or retrieval telemetry. A target removed by the score-margin stage is
reported with the closed-set reason margin_pruned; its detail includes the
observed relative margin (formatted to six decimals), the configured threshold,
and the retained top-row id. The existing omit, selective, scope, validity,
and budget reasons remain unchanged for rows rejected by those stages.
Change-intent lineage unfold — explicit recall only (issue #82)
Already-stored lineage (update_of from update) becomes visible on the
EXPLICIT path: when the query reads as change-intent ("what changed about X",
"why did we switch to Y", "we used to Z") AND the presented results contain a
live head of a tombstoned predecessor, recall appends the predecessor chain as
budgeted extra rows tagged [PREVIOUSLY] (JSON keys unfold_of +
unfold_hop, 1 = immediate predecessor). Extras never count against
--limit, never enter the telemetry bump set (popularity rewards query
matches, not neighbors), never cross namespaces, and are never silently
dropped for taint — the explicit surface prefixes [INJECTION RISK] /
[UNTRUSTED WEB] like any other row.
It runs ONLY when all of these hold: not --no-bump (hooks, PreCompact,
Hermes prefetch, and eval defaults never unfold), not --no-unfold, the
query matches the compiled change-intent regexes (deterministic, no LLM), and
link_hops >= 1 (the same contract that keeps search — CLI, MCP, and
Hermes _tool_search — byte-identical; MCP recall unfolds for free).
--explain describes lineage in verdict detail but never injects rows.
| Env var | Default | Meaning |
|---|---|---|
ZMEM_UNFOLD_TOP_K |
3 |
max presented hits to walk backward from (clamped 1-10) |
ZMEM_UNFOLD_MAX_HOPS |
3 |
max update_of hops per chain (clamped 1-10) |
ZMEM_UNFOLD_BUDGET |
4 |
hard cap on total [PREVIOUSLY] extras per recall (clamped 1-20) |
ZMEM_FAILURES_DB_TIMEOUT_S |
1.0 |
Stop-hook db reader busy-wait budget in seconds (clamped 0.1-5.0; invalid → default; failures --db-timeout overrides) |
ZMEM_ZCODE_DB |
~/.zcode/cli/db/db.sqlite |
ZCode episodic db the Stop-hook detector reads (empty/unset = default) |
ZMEM_REFLECT |
enabled | set to exactly 0 to disable the Stop hook entirely |
Confidence floors (issue #58, 3.8)
Four distinct floors live on the recall path. Each reflects a different
surface's precision-vs-coverage tradeoff. They are env-overridable; the
constants live in schema_meta.py.
| Constant | Default | Env override | Used by |
|---|---|---|---|
INJECT_FLOOR_PROMPT_DEFAULT |
0.25 | ZMEM_INJECT_FLOOR_PROMPT |
recall (UserPromptSubmit / PreCompact). Hard floor on FTS/vec results — anything below is dropped before scoring. |
INJECT_FLOOR_RECENT_DEFAULT |
0.5 | ZMEM_INJECT_FLOOR_RECENT |
recent (SessionStart / subagent recall). Tighter because the surface is high-confidence recent material, not query-best match. |
INJECT_FLOOR_GATE_NONE_DEFAULT |
0.4 | ZMEM_INJECT_FLOOR_GATE_NONE |
Hook selective-inject gate. signal=none rows must clear this floor; grounded-signal rows (test/compile/lint/reviewer/user) keep the 0.25 floor. |
INJECT_FLOOR_TRUST_DEFAULT |
0.2 | ZMEM_INJECT_FLOOR_TRUST |
Selective-inject gate (issue #115). Hard floor on trust_score: a row the contradiction ledger has driven below 0.2 (nine or more distinct contradict events) can no longer ride the passive lane; compute_score also multiplies every composite by trust_score (identity at the default 1.0, so uncontradicted rankings do not move — explicit recall/search keep retrieving the row, just ranked lower). When E-5 (#124) starts recording violations automatically, they must feed THIS ledger via trust deltas — violated_count never becomes a second independent gate input. |
The floors are intentional. Do not silently unify them. The
selective-inject gate (3.8) is a passive-lane filter; the trust floor
applies symmetrically to link-expansion neighbors that ride the passive
lane (a once-contradicted neighbor at trust 0.9 still renders with its
[CONTESTED LINK] marker).
Score-margin gate (issue #182)
ZMEM_INJECT_MARGIN is an opt-in score-separation gate between the existing
selective-inject filter and token-budget admission. It defaults to 0.0,
which disables the gate and preserves the legacy injection bytes. The setting
is read dynamically for each injection decision. A missing, negative,
non-numeric, NaN, or infinite value fails open to 0.0; values above 1.0
are clamped to 1.0. The recommended initial rollout is 0.05.
With a positive threshold and at least two usable post-selective rows, the
gate compares the two highest scores using the stable score-descending view
and computes the raw (top - second) / top. Every candidate score must be
usable; a malformed, scoreless, or non-finite score anywhere fails open. When
the raw margin is strictly below the threshold, ordinary rows after the winner
are pruned before the token budget runs. A decision or constraint in either
leading position is protected: no rows are pruned, but the valid observed
margin is still reported. Fewer than two rows and a non-positive top score
also fail open. The original presentation order is retained for surviving
rows.
When a valid score pair was observed, injection JSON adds margin formatted
to six decimal places and margin_pruned_ids in prune order; no fields are
added when the gate could not make a decision. The hook decision line appends
the same diagnostics, in the order margin= then margin_pruned=[...], after
the existing optional fields. Empty pruned lists are omitted from the hook
line, and malformed optional values are ignored independently so legacy
logging remains fail-open and byte-compatible.
When a passive inject surfaces nothing, it names WHICH gate fired (issue #87 /
#85 direction 1): no durable memories retrieved for this prompt. means the
candidate pool was empty OR every retrieved row was dropped by the passive
injection-risk filter — the one-liner is shared by design (the model is never
taught that omitted rows existed); the zmem-bg.log line distinguishes them.
no durable memories met the inject bar. means rows reached the
selective-inject gate and none passed. A token-budget wipe says so in its own
words (memories withheld: the injection token budget (ZMEM_INJECT_TOKEN_BUDGET) dropped every candidate row.). Hermes/MCP session_start use the session
variants (no durable memories retrieved for this session. and
session memories withheld: ...).
zmem-bg.log carries the same cut per decision line: every
zmem-hook line has reason= (reason=empty-pool, reason=omitted,
reason=below-bar, reason=budget-drop, reason=below-relevance,
reason=already-delivered, reason=injected), plus omitted=N
when the passive injection-risk filter dropped rows. The closed set lives in
schema_meta.py (INJECT_SILENT_REASONS). Since issue #94 every line also
ends with sid=<sanitized session id> ([^A-Za-z0-9._-] → _, cap 128;
sid=unknown when the host sent none) — the session key doctor's
--miss-rate join binds failures to injections with. The order is status, reason, omitted=, ids=, all=, tokens=, rendered_estimate=, admission_budget=, budget_dropped=, budget_truncated=, budget_dropped_protected=, sid= (the retired ops= ring count rode before sid= in releases <= 0.37; issue #158 moved operation context into the query, not the log). Current lines append exc=, moment=, optional lane=, ver=, t_ms=, then arms=, batch=, tools=, paths=, and margin fields (the budget fields are issue #116's distinct labeled numbers and ride only when budget accounting ran).
Decision attribution and report projection (issue #153)
The additive decision suffix is ordered lane=, ver=, t_ms= after the
historical moment= field and before later additive tails (arms=,
batch=, tools=, paths=, and margin fields). Closed lanes are claude,
codex, zcode, hermes-provider, and hermes-compat; runtime moments are
session_start, user_prompt, pretool, subagent, and precompact.
Silent reasons are exactly empty-pool, omitted, below-bar, budget-drop,
below-relevance, already-delivered, and expired (expired is reserved
for the expiry workstream and has no producer here). reason=injected is a
successful decision and reason=disabled is the separate kill-switch marker.
already-delivered is selected when pre-ledger candidate ids are nonempty but
the post-ledger pool is empty, after budget-drop in the precedence order;
dual-empty remains empty-pool. ver= and t_ms= are an all-or-nothing
enrichment pair: ver= is the release-manifest semver and t_ms= is a
nonnegative rounded perf_counter duration. A manifest/version failure keeps
the complete legacy line and omits all attribution fields. A lane is optional
for compatibility callers, but an explicit lane must be closed-set.
parse_bg_log accepts legacy lines without attribution. If any attribution
token appears, it requires valid ver= and decimal nonnegative t_ms= and
refuses malformed enrichment; moment= remains open for compatibility values.
Lane-less enriched lines and moments outside the report set stay aggregate-only.
The named report is always a sorted, zero-filled 20-row matrix: five lanes ×
the four report moments session_start, user_prompt, pretool, and
precompact. The runtime subagent moment and session_start_compact are
intentionally excluded from that 5×4 projection but remain visible in
aggregate statistics. The local Hermes provider writes hermes-provider; the
remote compatibility path writes hermes-compat, and invalid explicit lanes
return structured status 2 before store work while absent lanes stay omitted.
Query context (prior-turn operation tokens) — issue #88 / #85 direction 2
Decision-point prompts are prose with zero lexical overlap with the
operation-adjacent lessons that matter; the retrieval signal lives in the
tool commands the session is executing. The PostToolUse hooks
(convention-capture on the coding hosts, post_tool_call on Hermes) append
each Edit/Write/Bash event to a per-session ring at
<data>/ops/<session>.log, storing ONLY the tool name plus allowlisted
tokens (git subcommand chains, test-runner verbs, edited-path basenames) —
never a raw command dump, stdout, or argument values that are bare words;
secret-shaped tokens (common credential prefixes like ghp_/sk-/xox?-)
are dropped as well. The ring is byte-capped: past 64 KB it is trimmed to
the newest 64 lines. The store-owned passive selector composes that ring only
for the pretool moment, with the ops tail occupying a fixed reserved slice
INSIDE the 500-char cap (never appended past it; the separator shares the
slice, so no token is severed). UserPromptSubmit and Hermes prefetch no
longer append a caller-owned ops tail.
ZMEM_QUERY_CONTEXT=0 is the lane kill switch — it stops BOTH composition
and ring collection (an operator disabling the lane expects no sidecar
writes). Operation derivation remains store-owned; no operation count crosses
the exact hook envelope. Explicit surfaces (recall --query, search) are unchanged.
Limitation (deliberate, see #88): this helps LATER turns only — the first
tool call of a turn still runs before any operation context exists.
Cost note (#93 B4): credential-prefix-shaped tokens (sk-, npm_, AKIA,
…) are dropped from the ops query EVEN when they are legitimate filenames —
fail-safe direction (signal loss, never a leak); rename a real path that
collides. The eval composer ignores ZMEM_QUERY_CONTEXT by design (#93 B6):
evals must stay deterministic and immune to ambient env, so the kill switch
does not change their queries.
Pre-tool checkpoint phrases — issue #99
PreToolUse passive recall also enriches the query from the current raw
tool_input object before the tool runs. The store serializes that object as
sorted compact JSON (sort_keys=True, separators , and :), examines
exactly its first 150 Unicode characters, and selects at most one entry from
this immutable ordered table:
| Event | Checkpoint phrase |
|---|---|
| stash-consume | foreign-stash conflict verify stash list |
| reset | stale tree fetch main rebase verify diff |
| force-push | stale tree fetched base force-with-lease |
| branch-publication | stale tree fetched base force-with-lease |
| base-rewrite | base drift citation re-pin |
| path-test | basename ratchet citation re-pin local battery |
The recognized forms are git stash pop|apply|drop, git reset --soft|--hard,
git push --force|--force-with-lease|-f, ordinary git push, git merge --squash, and test/citation paths. git stash list, unknown operations,
non-object input, and matches beyond character 150 add no phrase. Current-event
operation tokens take precedence over an older session ring; the ring remains
the fallback. The operation-token tail retains its 150-character reservation,
the checkpoint phrase is never split, and the final query remains bounded to
500 Unicode characters. Raw tool input crosses only the private bounded
hook-to-store stdin channel and is neither logged nor persisted. The
ZMEM_QUERY_CONTEXT=0 kill switch disables both token and phrase enrichment.
This feature shipped under issue #155's published measurement decision (the
2026-09-19 fail-closed INSUFFICIENT outcome); #155's predeclared measurement
of record has since landed (eval/real-corpus-2026-09-19.json — an audit
over a frozen cohort, not a live efficacy claim), so the gate is satisfied
without an efficacy claim by this feature.
Passive-injection kill switch (ZMEM_INJECT=0) — issue #110 / P0-5
ZMEM_INJECT=0 disables every passive recall-injection surface: the shared
recall body (all modes — user_prompt, pretool, subagent, precompact/recent),
the SessionStart hook (Tier 2 recall AND Tier 0 — under the switch
SessionStart emits its empty {} envelope), the Hermes provider
(prefetch, the zmem_session_start tool twin, and the system-prompt
core.md block), the Hermes reflect hook's delivery paths, and MCP
session_start. Each silenced surface emits its empty envelope and logs
status=silent reason=disabled (reason=disabled is written only by this
switch, never by silent-reason classification — it lives beside injected
as INJECT_REASON_DISABLED, not in INJECT_SILENT_REASONS). Only the
literal 0 (whitespace-tolerated) disables — the ZMEM_QUERY_CONTEXT
convention; false/no/empty leave injection enabled. Capture paths never
consult the switch: correction capture, the ops ring, convention counters,
failure capture, and session-cadence maintenance keep writing — and the
capture-side PROMPT hooks (zmem-reflect.sh, zmem-subagent-reflect.sh,
zmem-convention-capture.sh, capture-failure/correction prompts) stay
active by design (#110: they prompt capture, they do not inject recalled
memory; the Hermes reflect hook is the documented divergence — a named
gated surface whose delivery carries recalled query-context). Parked
pre-tool fences and armed nudge markers are left in place and deliver on the
first enabled run. doctor shows the state (inject-switch line, WARN when
disabled). Near-miss spellings (0.0, 00, False, off) do NOT disable
— the parser matches the exact literal 0; verify with doctor after
setting.
Capture kill switch (ZMEM_CAPTURE=0) — issue #123
ZMEM_CAPTURE is a string environment variable with default "1". The
failure, convention, Stop, SubagentStop, and Hermes compatibility surfaces
read it before payload parsing, marker/queue writes, lesson queries, operation
appends, or other stateful work. Only a trimmed literal 0 disables capture;
undefined, empty, whitespace, false, and 00 stay enabled. Disabled shell
surfaces emit their exact empty sentinel and Hermes emits {}. Audits and
probes set ZMEM_CAPTURE=0 so those five covered surfaces cannot teach the
system.
This is independent of ZMEM_INJECT (delivery) and ZMEM_QUERY_CONTEXT
(operation-ring collection). Under ZMEM_INJECT=0, capture can still run;
under ZMEM_CAPTURE=0, no capture subprocess or state mutation runs on those
five surfaces. The pre-existing UserPromptSubmit correction queue and Hermes
convention-compatibility cadence are separate legacy paths with their own
capture-mode/interval controls. Complete
<<<ZMEM_UNTRUSTED_FENCE>>> through <<<END_ZMEM_UNTRUSTED_FENCE>>> blocks
are removed before transcript-derived content reaches deduplication, queue
synthesis, or closeout review.
Pre-tool inject — issue #90 / #85 direction C
On hosts whose pre-tool contract was probed and confirmed (ZCode: documented;
Claude: emitted — issue #117 retired the sidecar default in favor of the
per-session delivery ledger), a
PreToolUse hook (zmem-pretool-recall.sh, matcher
Edit|Write|MultiEdit|NotebookEdit|Bash|Agent since 0.30.0 — issue #119;
the Agent branch passes task text when present and otherwise uses the
store-owned recent path)
derives the recall query from the
tool input ITSELF — the command or file path about to run — and injects
matching hazard lessons before the tool executes. Pre-tool
additionalContext is documented on Claude Code (since 2.1.9 it lands
alongside the tool result; pausing is permissionDecision-driven only) —
issue #158 makes the store-owned selector the only passive path: delivered
ids now live in a bounded per-session ledger
(<data>/ops/<sha256-of-session-id>.ledger, atomic, window- and
cap-bounded) that EVERY injection moment consults, so the same row is not
re-delivered within the window — with one deliberate exception: PreToolUse
re-delivers when the operation tokens about to run strongly match the row
(the session-start hazard still fires before the dangerous command). The
ledger clears at PreCompact and the following SessionStart uses the ordinary
session_start selector moment. The retired ZMEM_PENDING_SIDECAR fallback
is no longer produced or consumed. The host event's session_id (without
it the adapter stays fail-open) is passed to the selector. The hook NEVER
denies (a surfaced hazard is information, not
grounds to block a legitimate command) and stays fully silent when nothing
qualified. ZMEM_QUERY_CONTEXT=0 silences every query-context lane, this
one included. Hermes delivers the equivalent on pre_llm_call (after the
fact of the producing call), best-effort per ring CURSOR
(ts, event-count) — same-second events still deliver; a transient store
failure after the at-most-once marker skips that cursor's delivery. Hermes
delivery's namespace follows the hook chain ZMEM_MCP_NAMESPACE →
ZMEM_NAMESPACE → user:global (issue #71 review: one chain for prefetch,
recall, and correction capture — Hermes hook events themselves carry no
namespace); project-scoped operation context delivers
on the coding-host PreToolUse surface. All query-context persistence
(rings, delivery markers, and the session delivery ledger) lives under <data>/ops/
sidecars and never grows the store's tables. Codex pre-tool injection is
WIRED (issue #95): hooks.codex.json registers PreToolUse with matcher
Bash|apply_patch — dumped live from codex-cli 0.153.0 (Windows,
2026-09-09: shell operations emit the hook tool name Bash, file patches
emit apply_patch; Codex treats an all-alphanumeric/pipe matcher string as
EXACT alternation, so it matches precisely those two names). MCP tools
(mcp__<server>__<tool>) and write_stdin are deliberately OUT — pre-tool
recall derives its query from shell/file-patch inputs, not arbitrary MCP
arguments. Upstream accepts hookSpecificOutput.additionalContext on
PreToolUse (model-visible, non-blocking); Codex hooks reference:
https://learn.chatgpt.com/docs/hooks. Codex envelopes are additionally
capped at 8000 encoded bytes (≈2000 tokens at the plugin's 4-chars/token
estimator; dense multi-byte (CJK) content has less real headroom) — 20% margin under upstream's 2,500-token hook-output spill
limit (DEFAULT_HOOK_OUTPUT_TOKEN_LIMIT, codex-rs output_spill.rs,
verified against tag rust-v0.153.0); the cap applies even when an operator
sets a larger ZMEM_CTX_BUDGET. Keep the verification-first convention: re-probe
with a live tool_name dump before changing the matcher.
Inject surface parity (host facts, not aspirations): Claude Code registers
SubagentStart (task-text recall — issue #119, probe 2026-09-10: neither
host's SubagentStart payload carries task text, so the query ladder is
payload-field-if-present → task/query text supplied by the event when
available → the queryless recent selector). The former delegating
PreToolUse(Agent) stash is retired; no zmem task-text sidecar is written or
consumed. Claude's matcher remains
Edit|Write|MultiEdit|NotebookEdit|Bash|Agent|Task since 0.31.0 — issue
#119; Task is accepted as the pre-rename delegation tool name per community
issue 29677, closed stale and not vendor-confirmed. Claude Code also registers
PostToolBatch (issue #120, 2026-09-11 — post-edit checkpoint recall, the
batch sibling of PreToolUse): after one completed batch, the hook parses
ONLY tool_uses[].name plus input.command / file_path /
notebook_path / path (plus the singular tool_name/tool_input
compatibility shape), bounds every retained field to 150 chars and the
joined query to 500, derives the existing operation tokens, and makes ONE
store.py recall --for-injection --no-bump decision with the #117 ledger
exclusion — raw tool_response/result payloads are never parsed, stored,
or logged (the decision log carries tool NAMES and path BASENAMES only).
The internal mode is posttoolbatch, but the runtime moment recorded in
the decision log and ledger is the closed-set pretool — no
posttoolbatch moment is emitted anywhere. PostToolBatch is deliberately
UNREGISTERED on Codex (no such upstream event exists in codex-rs hooks as
of the 0.153.0 probe) and is a documented host gap on ZCode (see the
seven-event note below); no lane reads tool_response — convention-capture
on PostToolUse parses the same tool_name/tool_input-shaped fields, not the
response payload. Codex registers SessionStart,
UserPromptSubmit,
PreToolUse (matcher Bash|apply_patch, probe 2026-09-09, codex-cli
0.153.0), PostToolUse, Stop, SubagentStart, SubagentStop, and PreCompact —
the Codex DELEGATION tool surface is UNVERIFIED by #119 (2026-09-10): the
#95 dump covered only the shell and patch/apply tools and cannot say
whether a delegation facility fires PreToolUse at all, so no matcher entry
was added and Codex SubagentStart rides the transcript-tail/recent rungs
(fresh delegation-tool dump deferred to the live-probe owner #96).
Upstream drops additionalContext on PreCompact (decision control only,
verified 2026-09-09 from codex-rs source), so zmem's Codex PreCompact
entry exists to clear the delivery ledger before compaction;
post-compaction re-injection rides
the registered SessionStart, which upstream fires with source=compact
after every compaction. PostCompact stays UNREGISTERED on Codex (issue
#118 decision, 2026-09-10: upstream Codex PostCompact carries only
trigger: manual|auto — no compact_summary, so there is nothing to
stash). The compact branch (issue #158): SessionStart branches on source == "compact" only — the launcher exports the payload field as
ZMEM_SESSION_SOURCE, clears the delivery ledger at PreCompact, and then
calls the store-owned selector with moment session_start on the next
SessionStart. The local decision log may retain
moment=session_start_compact for diagnostics, but that label is never sent
to storelib. There is no compact-summary sidecar or special second budget.
The payload block lives in
hooks/lib/zmem-session-start-payload.py — NEVER inline it back as
python -c: the string outgrew the Windows ~32K CreateProcess
command-line limit and silently degraded the hook to {}. Whether the PreCompact fence itself survives a live
/compact is UNPROBED — no claim either way until #96's live canary
lands (the host-capability rot convention from #103/#104). ZCode supports exactly
seven hook events — SessionStart, UserPromptSubmit, PreToolUse,
PermissionRequest, PostToolUse, PostToolUseFailure, Stop — so SubagentStart,
PreCompact, and PostCompact are host gaps on ZCode (an unsupported event name would be
dead config under the host's strict schema, so they are documented here
instead of registered; likewise ZCode's PreToolUse matcher deliberately
omits Agent — with no SubagentStart event, event-provided task text has no
consumer). If ZCode grows either event, wire
zmem-subagent-recall.sh / zmem-precompact.sh / zmem-postcompact.sh
immediately. PostToolBatch shares that ZCode gap (issue #120): until the
host grows the event, zmem-posttoolbatch-recall.sh stays
Claude-registered only, and the launcher's translation mapping for the
verb is inert on hosts whose manifest omits the entry; if ZCode grows the
event, wire zmem-posttoolbatch-recall.sh immediately (same convention).
Decision-point checkpoints (REQUIRED skill contract) — #85 direction E
When an agent is about to run one of the named hazardous operations, the skill/workflow driving it MUST first run an explicit recall and treat any hit as blocking review (read the lesson before proceeding; do not skip it because the task feels urgent):
- before
git stash pop/ any stash-consume —recall --query "git stash pop foreign stash conflict" --include-global(a blind pop can apply a foreign stash) - before
git reset --soft(squash assembly) —recall --query "git reset soft origin main stale tree" --include-global(a fetch may have moved the base) - before
git push—recall --query "git push stale tree fetch rebase verify" --include-global(verify the tree against the fetched base first) - before editing any file named by a stored citation/ratchet lesson —
recall --query "<path basename> ratchet citation re-pin" --include-global(cited-file edits have local gate batteries)
These checkpoints complement the pre-tool hook; they are not a substitute for it (an agent improvising a raw command is exactly who the hook covers).
Schema forward-compatibility (issue #65 follow-up)
A client refuses a store whose schema_version is above its ceiling.
FORWARD_COMPAT_SCHEMA_VERSION (schema_meta) is that ceiling: a maintenance
release of an older client line may raise it to an ADDITIVE-ONLY newer schema
(new side tables, no memory changes), letting the older client keep storing
and recalling memories — with a one-time stderr NOTICE — until it updates.
Above the ceiling the refusal names the update, and
ZMEM_ALLOW_NEWER_SCHEMA=1 overrides at the operator's own risk.
Injection token budget (issues #65 10.9, #116)
ZMEM_INJECT_TOKEN_BUDGET (default 1500, measured at 4 chars/token — no
tokenizer is bundled) is a HARD CEILING on the memories injected by the hooks
(UserPromptSubmit / PreCompact / SubagentStart / SessionStart Tier 2) and by
the session_start MCP/Hermes tools. Admission charges each row its FULL
rendered fence contribution plus the fence shell, with the same estimator the
reporting uses — the rendered fence never exceeds the budget. When the
ceiling is hit, admission skips past rows that do not fit (a later smaller
row is still admitted); decision/constraint rows are never dropped while
they can carry information: a row that fits renders whole, one that does not
is TRUNCATED with an explicit …[budget-truncated] marker, and only a row
too large for even a minimal stub is dropped — never silently. When anything
was omitted, the fence carries a machine-readable line
# [budget: dropped N rows, truncated M] and the shared hook body's
decision line (UserPromptSubmit / PreCompact / SubagentStart / SessionStart
tier-2 python path) reports tokens=<rendered>/<budget> plus the distinct
labeled fields rendered_estimate=, admission_budget=, budget_dropped=,
budget_truncated=, budget_dropped_protected= (the same order as the
field-order list above; emitted only when budget accounting ran). The
separate zmem-session-start.sh bash writer keeps its legacy
tokens=<used>/<budget>-only line. Every read --json envelope reports
tokens_used/tokens_budget and, on the --for-injection lane,
budget_dropped/budget_admission/budget_truncated/
budget_dropped_protected/budget_note. The B-1 report (doctor --miss-rate) counts
over-budget N decisions so the ceiling can be verified on a live log
window. ZMEM_CTX_BUDGET (character cap) remains the hard outer truncation
on the rendered block.
--include-global (opt-in) ALSO surfaces up to --global-limit query-relevant
rows from the user:global tier, merged project-first so a global row never
crowds out a project row. The three automatic hooks pass this so cross-project
lessons reach project-scoped sessions. Without it, behaviour is strict-namespaced
(byte-identical to before). When you want the global tier unioned in but still
want a per-tier budget, use recall --namespace project:<x> --include-global
rather than going unscoped.
Scoped five-tier recall (issue #167)
Ordinary implicit recall and recent calls without --namespace resolve the
current project, fleet, and host through the shared #166 scope resolver and use
five independent reservations in this order; hostnames are normalized to
lowercase for host: scope lookup:
project, domain, fleet_host, cross_project, user_global. The default
slot caps are 5/2/2/2/3. Library callers opt in explicitly with a scopes= map
on recall_memory, recent_memory, or the read-only explain_recall API; the
resolver's agent value is ignored. user_global is admitted only when
include_global=True.
ZMEM_TIER_SLOTS overrides the caps per call as exactly five comma-separated
ASCII nonnegative integers in that order. Empty, malformed, signed, or Unicode
numerals fail before the scoped pipeline opens a SQL lane. The scoped
cross_project slots remain reserved until a public admission policy is
available. While no policy is wired, recall and recent skip the unfiltered
candidate scan. This is separate from the legacy #98
--include-cross-project hazard lane. Scoped rows are stable-deduplicated by id
and carry their tier in JSON; generic fenced output prefixes [tier=<name>],
and a tierless generic row is explicitly prefixed [tier=unknown]. Ordinary
plain-text recall/recent output includes that marker too. Legacy passive
injection retains its established tierless wire bytes. The legacy
tier=cross renderer marker remains a suffix. Explicit --namespace, search,
hook, and injection calls retain their legacy route and limit arguments.
On implicit ordinary recall and recent, --include-global opts into the
scoped user_global reservation and preserves tier labels. Explicit
--namespace calls retain their legacy union behavior; --include-cross-project
retains its legacy hazard semantics. Direct scoped
recall_memory/recent_memory calls reject for_injection=True; scoped explain
remains read-only and accepts it. Programmatic callers use scopes= and
include_global=True for the same labeled global tier.
The domain reservation is available to programmatic callers that supply a
domain scope; shipped CLI, hook, and MCP resolvers do not currently create one.
Scoped recall/recent reject the legacy include_cross_project=True flag.
Cross-project hazard tier (issue #98)
A fourth, precision-gated tier (--include-cross-project, wired automatically
on the passive injection surface) can deliver up to 2 live rows from FOREIGN
project:* namespaces — a lesson another project already paid for, surfaced
exactly when you are about to repeat its incident. Admission requires ALL of:
- the running operation is hazardous: the derived ops tokens
(
derive_ops_tokens, the#88/#123allowlist) whole-token-intersect the hazard-verb set —ops_tokens._HAZARDOUS_SUBS(git push/reset/stash pop/ rebase/...) by default, overridable viaZMEM_CROSS_PROJECT_HAZARD_VERBS(comma-separated, trimmed, case-folded, de-duplicated; unknown or empty verbs are dropped with a one-shot stderr warning and an override with no usable verb falls back to the default set); - the row's
signalis grounded: one oftest,compile,lint,reviewer; - the row passes the standard score/confidence floor (#113) and is live
(
superseded_at IS NULL); - the row's namespace matches
project:*and is OUTSIDE the current project's alias set (user:globalrows stay in their own tier).
Surface policy (ZMEM_CROSS_PROJECT): unset → pretool only (PostToolBatch
maps to the pretool moment); 0 → off everywhere (wins even over an
explicit --include-cross-project); 1 → pretool and user_prompt (on an
env-enabled user_prompt surface the store-side selector derives ops tokens
from the prompt event itself — the #158 hook boundary keeps the hook a thin
flag forwarder); any other non-empty value → pretool only plus a one-shot
stderr warning. The tier is query-time — the queryless recent pull never
admits cross rows.
Scoped MCP tokens do not forward ops_tokens to this legacy path, so foreign
rows are excluded before rendering, context construction, or delivery-ledger
updates.
No-copy rule: a cross row is never copied, rewritten, or mirrored — the
store's own row renders in place, inside the untrusted fence, tagged
[ns=<source namespace>] [tier=cross] (with tier: "cross" on the JSON
row), and cross rows never consume project or global slots. Delivered cross
rows count once in surfaced_count under the same telemetry law as every
other tier. The #155 real-corpus replay baseline has landed: the lane ships
with these conservative defaults, and the predeclared measurement record
eval/real-corpus-2026-09-19.json (see the read-only replay evaluator section)
is the calibration reference — an audit over a frozen cohort, not a live
efficacy claim. recall --explain
does not include the cross tier (the read-only debugger predates it and is
not extended by #98).
Hybrid is the DEFAULT when embeddings are available (issue #58 3.3): the
query is embedded and matched against stored embeddings (sqlite-vec KNN),
then both lanes' rankings are fused with Reciprocal Rank Fusion (RRF, k=60).
--hybrid is kept as an explicit no-op alias for older invocations;
--no-hybrid forces the lexical (FTS5/BM25-only) lane. The embedding
runtime (onnxruntime + tokenizers, model lazy-downloaded and
checksum-verified) is optional and fails open: when it is unavailable,
recall silently uses plain keyword ranking — never an error.
Entity matching is the THIRD RRF list and needs no model (v10, issue #60
5.3): at recall the query is run through the same deterministic entity
extractor used on write, plus plain query tokens are matched against stored
entity aliases. Memories linked to a matched entity join the fusion as a
third ranked list (rank = number of matched entities, then recency). This is
how the alias rg recalls ripgrep-linked memories even when BM25 terms
miss — and it is live on model-absent stores by default. An unknown alias is
simply an empty third list; the lexical (and, when available, vector) lanes
fuse exactly as before.
MMR diversity is the DEFAULT (v10, issue #60 5.5): after RRF fusion and
composite scoring, the candidate set is re-ordered with Maximal Marginal
Relevance before --limit is applied, so a cluster of near-paraphrase
duplicates cannot crowd out a distinct fact. Similarity is embedding cosine
when both rows have embeddings, else Jaccard on normalized content tokens —
so MMR also works model-absent. --no-mmr returns pure composite-score
order (four paraphrases instead of three + the distinct fact). The tradeoff
knob is lambda: default 0.7, env ZMEM_MMR_LAMBDA (0.0 = maximize
Embedding profiles & rerank (ZMEM_EMBED_PROFILE, ZMEM_CROSS_ENCODER,
ZMEM_CROSS_ENCODER_MODEL): ZMEM_EMBED_PROFILE selects a row from the
registry in skills/memory/scripts/embed_profiles.py — unknown values refuse
with exit 2 before any store work, and a value whose dimension differs from
the store's committed vectors refuses until reembed --all converts it.
The cross-encoder rerank is enabled by default (issue #126 flip); set ZMEM_CROSS_ENCODER=0 to disable it. The
EXPLICIT lane covers CLI recall invocations that are not --no-bump/
--no-hybrid (never recent / search-aliases); pointing
ZMEM_CROSS_ENCODER_MODEL at a LOCAL pair-scoring .onnx plus sibling
tokenizer.json keeps the operator-vouched load path. Issue #125 adds a
checked-in mini-pair-scorer profile
(skills/memory/scripts/cross_encoder_profiles.py, exact SHA-256 pin) used
when ZMEM_CROSS_ENCODER_MODEL is unset — resolved from ZMEM_MODELS_DIR
or the shared models dir, and loaded ONLY after verify_profile_file
confirms the digest; ZMEM_CROSS_ENCODER_MODEL_URL plus exact
ZMEM_MODEL_AUTODOWNLOAD=1 permits one digest-verified, atomic-download
attempt at scorer-load time. A second opt-in, ZMEM_CROSS_ENCODER_PASSIVE=1,
lets the PASSIVE injection lane score its final admitted set through
rerank_final_injection_set (issue #125) in shadow-only mode:
ZMEM_CROSS_ENCODER_SHADOW=1 logs bounded rank deltas to
${ZMEM_DATA}/cross-encoder-shadow.jsonl and the lane's order is unchanged
until the recorded promotion gate (PASSIVE_PROMOTION_GATE:
#111 precision > 0.8978333333333333, p95 <= 250 ms, #129 false-injection <= 0.0) is
measured and flipped. Every rerank attempt is budgeted (250 ms default,
ZMEM_CROSS_ENCODER_BUDGET_MS) and emits exactly one terminal
[zmem] cross-encoder reason=<...> line on stderr; any failure degrades to
un-reranked results and rerank can never fail a recall. There is still NO
unverified-load escape hatch (ZMEM_MODEL_ALLOW_UNVERIFIED does not
exist).
diversity, 1.0 = no diversity — identical ordering to --no-mmr).
--no-mmr and --no-hybrid are independent flags and can be combined.
Recall rows carry entity cards: recall --json rows include
entities: [{id, kind, name}], and the fenced hook render shows at most
three entity NAMES per row (never ids). get --id shows the same links.
1-hop link expansion is the DEFAULT (v11, issue #61 6.3): after MMR,
each recalled memory's related/supports links are walked one hop and
up to --link-budget (default 2) extra neighbor rows are appended — a
memory that never matched the query but is associative-neighbor of a
match still surfaces. contradicts neighbors are included only when
they survive the confidence floor and are tagged [CONTESTED LINK] in
the fenced render (contested_link: true in JSON). Expansion rows carry
link_relation / link_of / link_score / contested_link keys; rows
that matched the query themselves never carry them, and expansion rows
never advance telemetry (popularity rewards query matches, not link
neighbors — so expansion cannot distort later rankings). --link-hops 0
disables the walk; --link-budget 0 is equivalent. search and recent
never expand.
--no-bump (opt-in) makes a recall passive: it records a surface event
(surfaced_count/last_surfaced) instead of advancing retrieval_count/
last_retrieved. The explicit-vs-passive split is the system contract: the
three automatic hooks and the Hermes provider prefetch are passive surfaces
(a background injection is not a usefulness signal); explicit recall — this
CLI without --no-bump, the MCP recall/search tools, and the Hermes
MemoryProvider search tool — intentionally bumps retrieval_count, because
an explicit read IS evidence the memory was useful (issue #21).
Issue #114 sharpened that contract in two ways. First, the automatic hooks and
SessionStart now pass --for-injection (with --no-bump): the selective
inject gate and the token budget run INSIDE that one store call, the returned
rows are exactly the rendered set, and surfaced_count advances only for
rendered QUERY-MATCHED rows (link neighbors render but never count; unfold
is explicit-recall-only and never runs on this lane) —
one subprocess, one decision, no second ack process. Second, ranking
popularity now reads retrieval_count ONLY: passive surfaces are still
recorded (promote/prune/consolidate consume them) but they no longer feed the
composite score, so passive pulls cannot inflate their own ranking (the
live-store shape surfaced_count=371, retrieval_count=0 was this loop). The
weight itself is unchanged; #124 will repoint it at applied/violated counters.
Store-owned passive selector and envelope — issue #158
All passive consumers use the single storelib entry point
select_and_budget_for_injection. It owns selection, delivery-ledger
exclusion and recording, pre-tool operation-token composition, the hard token
budget, and the canonical untrusted fenced rendering. The recall hooks,
SessionStart, and Hermes provider are subprocess-only adapters: they consume
only the envelope's string rendered field and never read SQLite, the
operation ring, a delivery ledger, or memory rows themselves. The selector's
closed moments are session_start, user_prompt, pretool, subagent, and
precompact; its lanes are claude, codex, zcode, hermes-provider, and
hermes-compat. A compact restart is logged locally as
session_start_compact, but the selector and ledger always use
session_start.
The exact session-aware CLI forms are:
python <store.py> recall --query "<text>" --for-injection --json \
--session-id=<id> --moment <session_start|user_prompt|pretool|subagent|precompact> \
--lane <claude|codex|zcode|hermes-provider|hermes-compat> \
[--ops-token <token>]...
python <store.py> recent --for-injection --json \
--session-id=<id> --moment <session_start|user_prompt|pretool|subagent|precompact> \
--lane <claude|codex|zcode|hermes-provider|hermes-compat> \
[--ops-token <token>]...
python <store.py> ledger-clear --session-id=<id>
python <store.py> delivery-clear --session-id=<id>
The passive recall and recent commands accept the additive attribution
flags --session-id=<id>, --moment, --lane, and repeatable --ops-token when
called with --for-injection --json. An empty query dispatches to recent
selection. An omitted --ops-token list lets the store read the pre-tool ring;
the ring is composed only for pretool, not for UserPromptSubmit or other
moments. For passive recall and recent, a session id requires a moment, and
a moment requires a session id. The equals form keeps session ids beginning
with - unambiguous. ledger-clear --session-id=<id> clears one
session's delivery ledger without opening SQLite. delivery-clear --session-id=<id> clears the ledger and legacy pending sidecar for SessionEnd
cleanup, also without opening SQLite; both commands are idempotent when their
target sidecars are absent.
The old hook-owned pending, compact-summary, and task-text sidecars are no
longer produced or consumed by these adapters. The D-2 #118 compaction-snapshot
handoff is documented above. This reflects the ownership history: issue #117 superseded the pending sidecar
with the per-session delivery ledger, and #158 completes that retirement for the passive adapters.
ZMEM_PENDING_SIDECAR=1
remains only as a documented no-op for operators upgrading from the
compatibility path. Precompact clears the ledger;
the next SessionStart performs ordinary session_start selection. The MCP
server's passive surface (prefetch, the reworked session_start below, and
the Hermes twins) now rides the same store-owned selector. This ownership move is schema-free: neither schema-version
constant changes.
Query-aware passive prefetch — issue #159
prefetch is the query-aware passive lane: one
select_and_budget_for_injection call and one complete JSON envelope — JSON
is the command's only output mode, and there is no consumer-side renderer or
second budget run.
python <store.py> prefetch --query "<text>" --namespace <ns> \
--session-id <id> --moment <session_start|user_prompt|pretool|subagent|precompact> \
[--lane <claude|codex|zcode|hermes-provider|hermes-compat>] \
[--ops-token <token>]... [--exclude <memory-id>]...
--query, --namespace, --session-id, and --moment are required;
--exclude (memory ids the caller already holds) is repeatable.
--for-injection, --no-bump, and --json are accepted as explicit markers
of the (only) passive mode the selector pins. The selector owns the
relevance/trust gate and the 1,500-token budget; prefetch never advances
retrieval_count (only surfaced_count and the delivery ledger may move).
Delivery is session-attributed: a second turn for the same session whose
ledger already holds the candidate rows returns the silent
already-delivered envelope; ledger-clear --session-id resets the ledger,
while SessionEnd uses delivery-clear --session-id to retire both sidecars.
The MCP server exposes the same one-call envelope as the prefetch tool
(query, namespace, session_id, moment required; optional lane —
validated against the five-value tuple, never defaulted to a host lane — and
ops_tokens), enforcing namespace scope and the ZMEM_INJECT=0 kill switch
before any store subprocess, and returning the complete selector envelope
plus the additive context alias equal to rendered. The MCP
server does not forward ops_tokens for scoped tokens: the legacy cross-project
admission path has no namespace-allow-list awareness, so scoped prefetch cannot
admit foreign project rows before rendering or ledger updates. The MCP
session_start tool rides the same store-owned queryless selector path via
recent --for-injection --json --session-id <id> --moment session_start,
returning that envelope (additive context alias; result/namespace/
ids retained as back-compat aliases); an omitted session_id falls back
to the legacy for-injection envelope without the ledger key.
add — capture a memory
python <store.py> add \
--namespace "project:<basename>" \
--type <fact|lesson|convention|preference|decision|constraint> \
--content "<the knowledge, specific and actionable>" \
--tags "comma,separated" \
--signal <test|compile|lint|reviewer|user|none> \
[--taint <trusted_internal|untrusted_tool|untrusted_web>] \
[--source-ref "file:<path>" | "session:<id>" | "
*Truncated - read the full file at https://github.com/zaxbysauce/zmem/blob/791e1de5f5420d7c66d0ef78b1bc67f22717f30d/skills/memory/SKILL.md.*
