Imported from TulipFarm/tulipfarm (
apps/eval/AGENTS.md). Install upstream withnpx skills add TulipFarm/tulipfarm --skill eval. Copyright stays with the author.
apps/eval
Offline eval harness. Runs a versioned Corpus of Eval Cases through the real Context assembler and the real Tool loop, scores deterministic Expectations, and prints a Scorecard.
Manually triggered before a release. Never runs in the product; nothing here is reachable at runtime.
Every noun above is defined in metadata/terminologies.md → Offline eval,
which is binding. Two words are easy to get wrong: an eval check is an Expectation, never an
"assertion" (Memory owns that), and a Sweep is never an "eval run" (a Trial contains real Runs).
Read on / Skip
Read on if your task touches: eval Cases, Expectation kinds, the Corpus hash, Sweep execution, Scorecard output, or wiring a new model binding.
Skip if you are changing the loop, the assembler or a Tool — this app consumes them unchanged. Only its own Cases need updating when their observable behaviour moves.
Map
| Path | Owns |
|---|---|
src/case.ts |
EvalCase, the Expectation union, loop limits. The vocabulary everything else reads. |
src/docx-fixture.ts |
Independent synthetic OOXML bytes; Case content follows a bounded paragraph prefix. |
CaseAttachment.xlsx · .pptx |
Shared independent Office fixtures: grounded content follows 200+ rows or lives only in speaker notes; real L3 provider-prompt Expectations must fail when extraction drops it. |
src/corpus.ts |
Loading and validating corpus/*.json; corpusHash — the content hash. |
src/scorer.ts |
scoreCase — pure and total. No I/O, no model, no clock. |
src/runner.ts |
runSweep and ModelBinding — the single seam this framework adds. |
src/matrix.ts |
runMatrix — the same Corpus across several models, one Scorecard each. |
src/bindings.ts |
resolveBindings — turns --model sonnet,terra into the bindings to measure. |
src/scripted.ts |
scriptedBinding — replays each Case's script. Free, deterministic. |
src/dispatch.ts |
Scripted results; an L2 hosting fixture uses the shipped repository-push contract and real Tool gate, never a scripted ownership verdict. |
src/model.ts |
PINNED_MODELS and pinnedBinding — the real-vendor binding. The only file here that touches a credential. |
src/retry.ts |
withRetry — transient vendor failures, retried and counted. |
src/spend.ts |
Spend totals: tokens, dollars, and what could not be priced. |
soul/ |
The Eval Soul: the frozen fixture business every Case is measured against. Ordinary tracked files. |
src/eval-soul.ts |
loadEvalSoul — copies the fixture to a throwaway git repo and reads it with the real SoulLoader; soulContext maps an Agent into the assembler. |
src/guardrails.ts |
Runs the Eval Soul's policy through production GuardrailsService and TurnGuardrails; safetyHostingAuthority selects offline platform minimums, never runtime identity trust. |
src/l3/ |
Persisted Chat and Routine tiers on in-process PGlite; MCP uses the production domain, account authority, Tool host, approval wait, Soul writer, publisher, and verified active-bundle loader. |
src/l3/mcp.ts |
MCP fixtures use tool-host's production prepareMcpToolCall; never reconstruct approval intents in eval. |
src/l3/mcp-accounts.ts, src/l3/mcp-provider.ts |
Seed account state through the real repository; replace only the external MCP provider, never account policy or Soul writes. |
src/l3/native-routine.ts |
Native Routine admission through the shared production service path, real invocation gateway, account authority, verified Soul bundle, and transactional inbox/Run/Artifact stores. |
src/l3/tier.ts |
integrationReply Cases apply production reply settlement after a real completed Chat Turn; provider transports remain API test scope. |
src/l3/soul-write.ts |
The soul_write Tool, over the real writer and the real publisher; definitionMode: plan first uses the production YAML Plan compiler. |
src/l3/resource-records.ts |
Journey-persistent in-memory Record repositories; Resource authoring reloads through the real loader and mutations use @tulipfarm/resources. |
src/l3/file-store.ts |
The one place file_create runs for real, so a Case can observe Chat draft versus saved File lifecycle and audience. |
src/l3/attachments.ts |
Shared extraction, typed Office refusals, and real Turn-driver screening over Case bytes. |
CaseAttachment.pdf · src/l3/pdf-*.test.ts |
Real PDF fixtures and visual accounting; readable fixtures execute file_read before any declared defective-version replacement. |
src/verdict.ts |
caseVerdict, scoreable — one Case collapsed into one word. Shared so the grid and a Baseline delta can never disagree. |
src/baseline.ts |
compareToBaseline — pure. Refuses a delta across two Corpora or two models. |
src/artifact.ts |
ScorecardArtifact read/write, harnessVersion, baselinePath. The durable form. |
src/release.ts |
applyBaseline — archive, compare, promote. The only place a Baseline is written. |
src/scorecard.ts |
Text rendering, including renderDelta. |
src/progress.ts |
progressReporter — live per-Trial output while a real Sweep runs. |
src/args.ts |
Option parsing. Split out so a mistyped ceiling is a tested refusal. |
src/cli.ts |
pnpm eval. |
baselines/<model>.json |
The promoted reference per model, committed so git history records the harness improving. |
corpus/ |
The Cases. One JSON file each, id matching the filename. |
Rules
The reasoning behind each of these — and thirty more — is in README.md. Read it
before changing how a Case is scored, run or compared. These are the ones that bite hardest:
- The Eval Soul is loaded, never constructed, and its hash is folded into
corpusHash. - The assembler stays in the path.
runTrialcallsassembleSystemPromptitself; a runner feeding hand-written prompts to the loop would never catch a Context-assembly regression. - Expectations are data, never functions — that is what makes the Corpus hashable.
- A vendor failure is not a verdict.
model_*loop failures count aserrored, neverfailed. corpusHashis the unit of comparability. A delta across two Corpora, or two models, is refused rather than computed.- Nothing becomes the Baseline on its own. Promotion is
--promote, from a clean tree, from a complete and unregressed Sweep only. - A content Expectation must be grounded in what the model was given, or declare
ungrounded. A File's declaredcontentcounts as given; its id and filename do not. attachmentsis what this Turn sent;readableis whatfile_readcan go and get. A File in both would be in the prompt from the first step, so a re-read Case assertingprompt_attacheswould pass with the whole mechanism removed. Declaring it in neither is refused for the same reason.- Name a shipped Tool with
platformTools; only hand-declare a Tool an operator would author. A copy measures the model against a description no deployment sends, and cannot assert the properties that live in the declaration. The resolved declaration is incorpusHash, so rewording one retires every Baseline. An unresolvable name fails the load. - Stateful L3 Tools run for real, routed by name in
routeTools. MCP setup/get and reviewed Tools use the production domain and live account authority; Soul/Resource authoring, Record mutations, andfile_createuse the real writer, loader, Resource service, and File store. Never script their successful result. MCP fixture account choices and consent are host state, not model arguments; the Corpus rejects scripted MCP results in every Turn. - MCP
revokeGrantBeforeApprovalremoves a durable grant while the real approval wait is pending. MCP reads must usemcpIntegrationsFromBundle, never authored-loader state. mcp_provider_call_countcounts actual externaltools/callrequests, not dispatcher attempts or model claims. Missing observations fail even when the expected count is zero.run_event_emitted.countpins an exact positive event count; repeated approvals must emit distinct durable events, not merely complete two provider calls.nativeRoutinestarts from a persisted, claimed native inbox event, not an HTTP webhook.native_admission_equalsobserves admission rows and real account resolution, never model text. Its lease fault changes the inbox fence after Run insertion to prove the shared transaction rolls back.- A saved generated File's audience comes from
agentRoles, never the soul'sroles:, which is advisory and writes norole_assignments. Pairgenerated_file_readable_bywith agenerated_file_not_readable_byfor a Role the Agent lacks, or the Case passes just as well against an audience that shares every File with everyone. - Every Case carries a
script, so the whole Corpus runs free in ordinary CI. - L3 preserves model streaming.
run_event_text_omitsreads concatenated durable participant text, not the final Message or operator evidence; missing observations fail. - L3 claims each Run before execution, so State writes use the real lease-generation fence.
- Both tiers collect guard refusals. L3 reads durable
guardrail.decisionevents. Use stage-specificguardrail_blockedso safe model refusals can be reported as unexercised. - Two models are a control, not a contest, and they never run at once.
- A Judge failure errors the Trial; it never scores low.
- A
faultis L3-only and breaks a dependency before or during a Turn. The checkpoint fault resumes a real approval checkpoint before failing; the output-limit fault uses the production model completion guard, not a synthetic error. - Persisted Message Expectations read
eval_messages, never Run events or recomputed history. Failure Cases must create their evidence through scripted model and Tool rounds before the fault. - Provider File Expectations read
splitPromptbinary parts against immutable Case Files. provider_prompt_containsreads the firstsplitPromptprojection, excluding assistant Messages. DOCX fixtures require a single L3 Chat Turn; text fixtures ground their final paragraph incontent.model_not_calledrequires an observed zero invocation count, not missing usage. DOCX refusal variants contain real defective archives, never scripted parser verdicts or invented content.- PDF Cases compare original binary bytes and all page estimates. A readable PDF's
replaceAfterReadchanges fixture storage after the real Tool succeeds; never script that Tool's result. - Checkpoint replay is one L2 fault seam. It crashes after the first Tool result and retries the same input; do not turn it into a general workflow fixture.
- This workspace is CJS-by-default — no
import.meta.
Commands
| Command | Runs |
|---|---|
pnpm eval |
The capability Corpus on the free scripted tier. |
pnpm eval:redteam |
The attack Corpus, scripted. |
pnpm eval:sonnet / eval:terra |
One real seat. Needs that seat's credential. |
pnpm eval:matrix / eval:redteam:matrix |
Both seats, one Scorecard each plus a grid. |
pnpm eval -- --help |
Every flag. |