Imported from zhu1090093659/plotforge (
AGENTS.md). Install upstream withnpx skills add zhu1090093659/plotforge. Copyright stays with the author.
PlotForge Agent Rules
This file is the project-level memory for future agents. Keep durable engineering constraints here; do not use it as a task log.
Project Goal
Build PlotForge as a CLI-first Rust engine with a creator desktop UI adapter. The engine includes opt-in real text, image, TTS, and moderation provider paths plus MCP tools, while local-pi remains the default no-network text path. Steamworks upload, hosted sharing, and paid Workshop integrations remain deferred until their planned tasks.
The current MVP contains:
- Rust workspace crates for schema, storage, rule evaluation, story craft review, local and real-provider agent planning, MCP, runtime, media asset registry, job queue core, static export, Studio adapters, and CLI.
apps/player-web/staticas the source for the no-network static player package copied byplotforge-export.apps/creator-desktopas the Vite/React/Tailwind Studio frontend workspace with dashboard, source editor, playtest, runtime trace views, and a thin Tauri command bridge.- Folder project source of truth using TOML, JSON, Markdown, and generated local assets.
- Rebuildable SQLite cache/index under
.plotforge/cache.sqlitefor project summary, source file hashes, trace metadata, and asset metadata. - Export profiles and
ai-usage.jsonas redaction-safe package disclosure surfaces; they describe capabilities and AI usage evidence but do not provide legal conclusions or platform approval promises. plotforge-workshopas a local-only Workshop package schema and validator; it validates draft package metadata, file hashes, and disclosure files without Steam API or upload dependencies.- Steam Submission Kit draft generation in
plotforge-workshop; it emits local checklist, AI disclosure, content warning, and packaging-note Markdown drafts without legal, approval, or upload claims. - Steam compliance QA boundary in
docs/steam-compliance-qa.md, enforced byscripts/qa/no_launch_promise_lint.pyfor active project-facing guidance surfaces. - No default/example project is committed or auto-loaded; smoke flows create temporary starter projects with
plotforge new project. - Static web export served from local files over HTTP.
- Spec-driven GitHub issues are task/progress units, but pull requests are Phase-level delivery units; do not open one PR per small task during PRD completion work.
Source of Truth
plotforge-schemais the only schema contract source of truth.- Generated frontend contracts live in
contracts/; regenerate them withscripts/contracts/export_contracts.shafter Rust schema changes and verify withscripts/contracts/check_contracts.sh. - Folder project files are the source of truth for game projects; generated caches, exports, traces, and build artifacts must be rebuildable.
- Structured editing documents for World, StoryCraft, Characters, State variables, and Rules are schema-defined contracts backed by
plotforge-storagesource files; regenerate frontend contracts after changing them. - SQLite cache/index files are never required to load canonical project state; when cache contents conflict with folder files, folder files win and the cache must be rebuilt.
- Temporary starter projects generated by
plotforge new projectare the smoke-test source of truth for CLI/export flows. - Runtime traces are generated evidence, not committed fixture state.
- Runtime traces must use structured, redaction-safe fields for action intent, rule result, planner result, diagnostics, media references, fallback, and errors; do not write raw provider responses, secrets, or unredacted key markers into trace/debug output.
- Runtime traces, snapshots, and provider output envelopes must carry reproducibility metadata (
run_seed,prompt_version,model_version,provider_config_hash, optionalmcp_tool_call_hash/moderation_config_hash, and trace/snapshot evidence ids where applicable) without storing raw provider responses or credentials. - The retired
docs/design/ui-mockups/agent-native-v1/PNGs/README were historical development-time visual references only and are no longer present; the original design intent for the Agent Mesh, Command Center, Director Mode, proof, trace, and export package workflows now lives in the running Studio UI (apps/creator-desktop/src). Do not reintroduce those mockups and do not reference them fromapps/creator-desktop/src. - Real provider configuration is local-only: credentials must be injected through explicit local resolvers such as environment variables, provider config hashes must be derived only from non-secret config fields, and
providers/orprovider_config.*files must never become project source, contracts, traces, or export package content. - The user-global Provider registry (
~/.plotforge/providers.jsonor the platform config-dir equivalent), the Skills library (~/.plotforge/skills/,~/.plotforge/skill-index.json), and the MCP server registry (~/.plotforge/mcp.json) are user-global local-only configuration; they never enter project source, contracts, traces, or export packages. Credentials are referenced indirectly by environment-variable name (credential_env_var); the registries never store credential values. Every Provider registry mutation must use the agent-owned locked transaction boundary covering the complete load-modify-write cycle and an atomic same-directory replacement; Studio/Tauri/CLI adapters must not implement separate registry locks or direct read-modify-write sequences. - Prompt templates are split into a user-global library (
~/.plotforge/prompts.json, shared across projects) and an optional per-project store (<project>/.plotforge/prompts.json). Project-scoped templates stay under.plotforge/and are blocked from export packages by the export allowlist and the workshop denylist; user-global templates never enter any project. - Built-in prompt templates are code-managed, versioned compile-time constants (
PromptTemplate { id, version, system_message, output_instructions }inplotforge-agent::prompts), not user-edited files; theirversionstring (e.g.scene_planner_v1,beat_writer_v1,plot_doctor_v1) flows intoReproducibilityMetadata.prompt_versionon every agent output envelope so a change to a template's system/output text is a version bump that reproducibly distinguishes runs. Editing a built-in template's wording is a version increment, never a silent edit. - Real provider calls (text, image, TTS, and moderation) are in scope for
plotforge-agentand shipped behind the production HTTP clients (OpenAiCompatibleClient/OpenAiResponsesClient/AnthropicMessagesClientfor text,OpenAiImageClientfor images,OpenAiTtsClientfor TTS, andOpenAiModerationClientfor moderation). The offlinelocal-pipath remains the default no-network scene-planning path and derives its proposal from creator input and project context; generic fake providers are test-only. A fully no-network apply also leaves real media providers disabled. Real provider calls occur only through explicit enabled configuration. Provider configuration and recorded evidence stay local-only and redaction-safe: no raw provider response body, credential, or unredacted key marker enters traces, the persisted registry, project source, contracts, or export packages. - Dynamic model discovery (
fetch_provider_models) queries a provider's upstream/models(or Anthropic/v1/models) endpoint and caches the list locally under~/.plotforge/cache/models/{provider_id}.jsonwith a 1-hour TTL (MODEL_DISCOVERY_TTL_SECS = 3600); a fresh cache is returned as-is and a stale/missing cache triggers a fresh fetch. The cache stores only redaction-safeRemoteModelInfometadata (id,owned_by,created,max_input_tokens,max_output_tokens) — never credentials, endpoint response bodies, or secrets. The credential is resolved through the same strictEnvCredentialResolveras completions; a missing credential surfacesModelDiscoveryError::MissingCredential, never a silent empty list. - External Skill auto-discovery scans a fixed set of roots:
~/.plotforge/skills,~/.claude/skills,~/.codex/skills,~/.zcode/skills,~/.cursor/skills,~/.copilot/skills,~/.hanako/skills,~/.openclaw/skills,~/.workbuddy/skills,~/.redbox/skills, and~/.codex/vendor_imports/skills/skills/.curated. Roots that do not exist on this machine are silently skipped (not errors, not silent fallbacks). PlotForge never creates directories under any external root; it only reads and copies into its own~/.plotforge/skills/library. - Media asset records must use structured, redaction-safe provider metadata only; store prompt hashes/request ids when needed, never raw provider responses or secrets.
- Job records must use typed state, explicit failure objects, injected clocks for deterministic tests, and no hidden global async state.
- The user-global
UsageLedgerlives at~/.plotforge/usage.json(or the platform config-dir equivalent) and is owned byplotforge-job. It stores only redaction-safe per-call provider/model identifiers, generation kind, token counts, cost units, and timestamps; it never stores prompt content, response bodies, credentials, or secret markers, and it never enters project source, contracts, traces, or export packages. Missing ledger files mean no recorded usage; corrupt files fail explicitly rather than silently resetting accounting. Path-backed appends must reload under the ledger's cross-process lock and atomically replace the file in the same directory, so concurrent or stale handles cannot overwrite another call's usage; adapters usereport_usagerather than owning a second persistence path. Non-local apply flows validate the ledger before any moderation/text/MCP request, then share that ledger through moderation, text/MCP, and image accounting;local-pikeeps the no-ledger fast path. Every successful moderation response is recorded exactly once asUsageKind::Moderationbefore the flagged/pass branch, using stable provider ids and zero token counts when the upstream omits usage, so blocked calls and cost units remain visible. - Image provider fallbacks must remain trace-visible and must register placeholder assets as fallback metadata, not as successful generated-cache hits.
- Spec-driven planning artifacts and progress logs are local/private operator state for this open-source repository; keep them out of Git and under ignored paths unless explicitly approved.
- For broad PRD completion phases, prefer parallel sub-agents for disjoint implementation lanes after shared schema/contracts are planned; keep final integration, validation, and Phase PR scope decisions centralized.
Architecture Boundaries
- Keep rule evaluation in
plotforge-rule; do not duplicate rule behavior in CLI, UI, storage, export, or tests. - Keep runtime state transitions in
plotforge-runtime; agents propose content and runtime/rules commit state. - Runtime owns Scene/Beat progression: same-scene choices such as
continueadvance throughBeatNext::Beatwithout planner calls, scene changes, new images, or turn increments;change_scenechoices cross the planner/runtime boundary, set the next scene entry beat, and increment the turn. ProductionRuntimeSession::newusesNoScenePlanner, so crossing a scene boundary without pi-Agent or an explicitly injected provider fails withscene_planner_required;MockAgentPipelineis test-only. - Missing current beats, entry beats, or same-scene beat transitions must fail explicitly; do not silently fall back to another beat when committing runtime state.
- Player/freeform input must resolve to a typed
ActionIntent; unsupported input must not mutate runtime state or silently map to a default action. - Keep persistence and fixture file layout in
plotforge-storage. - Keep structured editing read/update/create behavior in
plotforge-storageand expose it throughplotforge-studio/Tauri command adapters; Creator Desktop must call the generated-contract bridge instead of parsing or validating project files in TypeScript. - Keep SQLite cache/index behavior in
plotforge-storage; it may index project summaries, source file hashes, trace metadata, and asset metadata, but must not become a second project loader or source of truth. - Keep static export behavior in
plotforge-export; exported bundles must copy only reachable referenced assets fromplotforge-mediaand must not include private traces, provider config, raw provider responses, unreferenced assets, or secrets. - Export profiles and AI usage manifests must stay schema-defined, redaction-safe, and capability/descriptive only; do not put provider credentials, raw provider responses, private traces, legal conclusions, or platform approval promises into them.
- Keep Workshop package validation in
plotforge-workshop; it must remain local/offline validation of draft package metadata and file hashes, not a Steamworks SDK wrapper, upload client, or release-readiness oracle. - Keep Steam Submission Kit generation in
plotforge-workshop; it may generate local draft documents from validated package evidence, but must not claim compliance, approval, publishing automation, or legal conclusions. - Keep Steam-facing docs and generated guidance under the no-launch-promise QA boundary; do not promise automatic publishing, platform outcomes, legal conclusions, or ownership of a creator's Steamworks workflow.
- Keep static player package behavior in
apps/player-web/static; it must consumeExportManifest, run without external network URLs, and must not duplicate rule/runtime state machines. The player is split intoplayer-core.jsplusplayer-save.js/player-i18n.js/player-audio.js/player-types.jsmodules; every module file must be listed in bothplotforge-export'sPLAYER_PACKAGE_FILESandscripts/qa/static_export_http_smoke.py's expected whitelist (defence-in-depth), and the smoke test must assert no network URLs across all player JS. - Keep asset registry records, content hashes, references, provider metadata, and reachability in
plotforge-media; as export, runtime, Studio, and future providers integrate media, consume this boundary instead of adding a second media registry implementation. - Keep long-running task state transitions, cancel/retry/timeout/progress, and cost accounting in
plotforge-job; provider/runtime/export adapters should not own parallel job state machines. - Keep the shared pre-flight
ThrottleGateinplotforge-job, configured through its provider-agnosticThrottleConfig.plotforge-agentowns one explicit process-level throttle registry keyed by provider kind, the stable registryid, and a redaction-safe hash of non-quota upstream identity (API shape/endpoint/model/credential env-var name), so provider adapters rebuilt between Studio turns reuse the same RPM/concurrency state without coupling unrelated registry contexts or entries that reuse an id; moderation normalizes its final request endpoint before hashing so equivalent URL spellings share rate debt. Quota edits reconfigure a gate in place and preserve in-flight permits and consumed rate capacity. Editable provider labels are display-only and must never identify throttle or usage-ledger records. With no quota fields, adapters keep the legacy fast path and do not load the ledger. Provider adapters acquire the gate before entering retry handling; the gate never retries, and the shared textRetryPolicyremains the single retry owner for both one-shot and MCP provider rounds. Rate limits surface an exactretry_after_ms, and concurrency permits release through RAII on every drop path.daily_token_budgetis text-only and is enforced from theUsageLedger's current UTC epoch-day output tokens; image, TTS, and moderation entries reject it because those upstream responses do not provide trustworthy output-token accounting.BudgetExceededis explicitly non-retryable. - Keep image provider ports and scene image pipeline orchestration in
plotforge-agent; image providers should feedplotforge-mediaandplotforge-jobinstead of bypassing those boundaries. - Keep CLI behavior in
plotforge-cli; CLI should orchestrate crates instead of owning business logic. CLI output formatting (i18n chrome terms, labels, colorized summaries, report rendering) lives inplotforge-cli/src/cli_output.rs;main.rsonly dispatches commands and orchestrates handlers. Color is emitted only when stdout is a TTY (std::io::IsTerminal), so pipes, redirects, and test harnesses see plain text. JSON output (studio commands,serde_json) is never colored and its field shape is a contract — never adjust JSON fields for human-readable changes. - CLI interactive wizards (e.g.
workshop submission-kit) usedialoguerto prompt step-by-step and are the default;--batchexplicitly enables full flag-driven input for scripts/tests. Wizards must guard stdin TTY (std::io::IsTerminal) and bail with an explicit error directing to--batchwhen non-TTY — never hang on a blocked stdin or silently fall back. Wizards only collect parameters; business validation and generation stay in the owning crate (plotforge-workshop), never reimplemented in the CLI. - Keep
apps/creator-desktopas an adapter over generated contracts and future Tauri commands. TypeScript UI code may import types fromcontracts/plotforge.d.ts, but must not reimplement rule, runtime, storage, storycraft, or agent business logic. - Creator Desktop UI is split into one View component per workspace region (
App.tsxrendersLaunchpadView[=Home],PlayView,WorldView,StoryView,CharactersView,StateView,RulesView,AssetMaintenanceView[=Assets],TraceDebugView[=Trace],ExportView,SourceView,SettingsView[=Settings]). Navigation is a single flat list (studioModel.ts: studioSections), not a two-level workflow/section tree.App.tsxonly orchestrates routing, layout, and form-action sinking intouseStudioWorkspacesub-hooks (useProjectEditing/usePlaytest/useExport/useAgentConversation). New regions go in their own*View.tsx, never as inlinerender*Panelfunctions inApp.tsx.SourceViewowns the project source-file browser and inline text editor;LaunchpadView(Home) is a project-entry + health surface only (no source editing, no run button, no director-input box) and renders insideStudioShellso the sidebar, agent rail, command palette, and collapse toggles are reachable from Home — Home is no longer a full-screen branch outside the shell. The right-sideAgentChatRailis the single "describe a change / run a turn" entry point — the sharedplaytestInputstate lives inusePlaytestand is consumed byuseAgentConversation; no other view may render a director-intent textarea or "run proof" button.PlayViewis the single read-only scene preview (background image + current beat + choices); clicking a choice submitschoice.labelas the next turn's intent viaagent.submit(). Snapshot controls (saveId/restoreId/restoreLatest) live in a "Advanced snapshot controls"CollapsibleinsideTraceDebugView, not in a separate Director Mode view. The honesty-surface boundary evidence lives on demand inside theAgentChatRail"Evidence" popover (aCollapsible"Local boundaries" block), not as an always-painted panel or a top-levelAgentMeshView. Structured-document creation (e.g.createCharacterFromDraft,createRuleFromDraft) must sink toplotforge-studioRust commands via the contract bridge, never be reimplemented in TypeScript.runtimeTraceView.tsxis an orphan skeleton superseded byTraceDebugView; do not revive it. The retiredCommandCenterView,DirectorModeView,AgentMeshView, andArtifactReviewViewwere deleted when the navigation was flattened; do not reintroduce them or their two-level workflow grouping. SettingsView(sidebar item #12) is the single entry point for Agent / MCP / Skill configuration: it owns three internal Tabs (Agent / MCP / Skills). The Agent tab (AgentConfigSection) covers text, image, TTS, and moderation provider subsections plus Model, Prompts, and the read-only usage summary; moderation remains a subsection here, never a fourth top-level Settings tab. The MCP tab (McpSection) renders the real MCP server management surface (list, add/edit/delete, stdio/sse/http transport config,Test connection→McpServerTestResult, per-project enable toggle persisted toAgentSessionConfig.enabled_mcp_servers, tool list withInvoke→ redactedMcpToolCallResult) — it is no longer a "coming in a future release" placeholder. The AGENTS.md boundary carve-out for MCP server calls has landed (see the MCP carve-out paragraph below); the schema types,plotforge-mcpcrate, Studio commands, contracts, and DataSource methods are implemented. The Skills tab (SkillsSection) covers the user skill library and per-project enablement. Provider entries carry onlycredential_env_var(the env-var name), never the credential value; the UI must never display or collect secret values. Model settings (model_id / permission / thinking / enabled_skills / enabled_mcp_servers) are persisted throughuseAgentConfig→set_agent_session_config. The sharedagentConfigSelectors.tsx(ModelSelect/PermissionSelect/ThinkingSelect) is the single source of truth for the model/permission/thinking selectors so the Home toolbar (LaunchpadView) and the Settings → Agent tab render the sameAgentSessionConfigwith the same UI. Prompt templates are split user-global / project-scoped. Skills are auto-discovered from external roots and imported into the user library; enabling a skill for a project writes its id toAgentSessionConfig.enabled_skills.- Keep HTTP provider client behavior in
plotforge-agent; the three supported API formats (openai_compatible/openai_responses/anthropic_messages) all implement theTextModelClienttrait, are constructed bybuild_provider_client/build_text_providerfrom aProviderEntry, and resolve credentials throughEnvCredentialResolver/OptionalEnvCredentialResolver. Credentials are never stored in struct fields, traces, or the persisted registry. Skills auto-discovery scans the fixed external roots (see Source of Truth) and stores only metadata inSkillManifest; skill bodies are loaded on demand viaload_skill_body(progressive disclosure). - Provider HTTP error handling is coordinated, not ad-hoc: a non-2xx
429 Too Many Requestssurfaces theRateLimit { retry_after_ms }kind (parsed from theRetry-Afterheader, seconds or HTTP-date), other non-2xx surface asProvider, and timeouts surface asTimeout. The text retry policy (RetryPolicy::default()) caps atmax_attempts = 3with exponential backoff (base_delay_ms = 500,max_delay_ms = 8_000, jittered) and honours the server-advisedretry_after_ms(clamped tomax_delay_ms) when present. Retryable kinds areRateLimit,Provider(5xx), andTimeout;ContentFilteredandOutputTruncatedare non-retryable (retrying with the same prompt/limit reproduces the same outcome and burns quota), so the retry loop surfaces them immediately with an attempt-count annotation in the message. Image and TTS providers mirror this:429→RateLimit, content-filter → non-retryableContentFiltered, with their ownIMAGE_JOB_MAX_ATTEMPTS = 2/TTS_JOB_MAX_ATTEMPTS = 2job-level caps. - Image and TTS providers are registered in their own
ProviderRegistryarrays (image_providers: Vec<ImageProviderEntry>,tts_providers: Vec<TtsProviderEntry>), separate from the textproviderslist, and resolved byresolve_image_provider/resolve_tts_provider(first enabled entry, orNone— the pi-Agent / TTS pipelines skip generation whenNonerather than failing the turn). Both entry types follow the samecredential_env_var-name pattern asProviderEntry: the registry stores the env-var name only, never the value, and the value is resolved at call time (OpenAiImageClientvia the injectedProviderCredentialResolver;OpenAiTtsClientviaEnvCredentialResolverinternally). A missing/empty credential for an auth-required provider surfaces an explicitimage_provider_missing_credential/tts_provider_missing_credentialerror — no silent degraded run.image_providersandtts_providersare#[serde(default)]so a registry serialized before they existed still deserializes. TextModelRequestcarries an additivemessages: Option<Vec<ChatMessage>>field (T2.1/T2.2). WhenSome, the three HTTP clients send the structured{role, content}conversation verbatim per their wire format (OpenAI Chat Completions →messagesarray; OpenAI Responses →instructions+inputsplit; Anthropic Messages → top-levelsystem+messagesarray); whenNone, the singlepromptis wrapped as one user message (the pre-T2.2 shape). This field is strictly additive: theTextModelClienttrait signature and the threecomplete()methods are NOT modified, andNonepreserves the prior behaviour. A secret-marker scan runs over each message's content at the same redaction boundary aspromptbefore it leaves the process.TextProviderConfigcarries asupports_json_schema: boolflag (T2.4). Whentrue, the HTTP clients use provider-native JSON Schema mode to constrain output to a validAgentOutputEnvelopeshape: OpenAI-compatible/Responses sendresponse_format: {"type":"json_schema","json_schema":{"name":"agent_output_envelope","schema":<AgentOutputEnvelope schema>,"strict":true}}; Anthropic Messages defines a forcedemit_envelopetool with the envelope as itsinput_schemaand parses thetool_useblock'sinput. Whenfalse, OpenAI-compatible keeps thejson_objectresponse format and Anthropic sends plain messages with no tool constraint (the pre-T2.4 shapes). The flag is deliberately excluded fromprovider_config_hash— it is a client-side capability flag, not part of provider identity. The schema is generated once viaschemars::schema_for!(AgentOutputEnvelope)and matches the sameJsonSchemaimpl that backscontracts/plotforge.schema.json.AgentChatRail's "describe a change / run a turn" drives thepi_agent_apply_runStudio command: the pi-Agent generates aScenePlanproposal via the configured provider (offline input-derived provider whenmodel_id == "local-pi", or a real HTTP provider),RuntimeSession::apply_agent_scene_planevaluates rules (change_sceneaction type) and commits the proposal throughcommit_scene_plan, and the result is appended as a chat turn carrying aPlayOnceReport-shapedPiAgentApplyResult.usePlaytestremains the owner of snapshot-control state (playtestSaveId/restoreId/restoreLatest) and the sharedplaytestInput; the rail mirrors the latest report intoplaytest.setPlaytestReportsoTraceDebugView/PlayViewstay in sync, but the rail's main submit path no longer drivesplay_once_project*. Failure modes are explicit (pi_agent_missing_credential,pi_agent_provider_timeout,pi_agent_unsupported_payload); no silent fallback.- Agent-native Creator Desktop UI may show pi-Agent runtime, agent capabilities, approvals, and evidence as explicit local workflow surfaces. Real internal pi-Agent execution (local, schema-backed, redaction-safe, wired through
plotforge-agent/plotforge-studiovia thepi_agent_run/pi_agent_capabilitiesStudio commands) is permitted and must stay local-only. External agent execution, hidden network model calls, Steam upload automation, legal conclusions, or platform approval remain forbidden unless those schema-backed integrations are explicitly added. - Keep i18n as an adapter concern for UI chrome and command output, with full English/Chinese coverage for every user-visible chrome string. Creator Desktop text is localized through
apps/creator-desktop/src/i18n.tsx; static player text is localized throughapps/player-web/static/player-i18n.js(split from player-core in T5.1); CLI output uses theOutputLanguageselector (--language/PLOTFORGE_LANGUAGE). Do not automatically translate project-authored titles, source files, story text, manifest content, or creator-provided export metadata. - Keep Tauri command behavior in testable Rust adapter crates such as
plotforge-studio;apps/creator-desktop/src-taurishould stay a thin IPC wrapper. - Do not add Steam/Workshop upload integrations, provider SDKs, or unplanned networked model calls to MVP core crates unless the project scope is explicitly changed. The real HTTP provider family in
plotforge-agent—text, image, TTS, and moderation—is explicitly in scope and is not a violation of this rule. These are first-partyreqwestclients built from schema-defined registry entries, never third-party provider SDKs; their configuration and recorded evidence stay local-only + redaction-safe per the Source of Truth and moderation carve-out rules. The rule still forbids pulling a vendor SDK (OpenAI/Anthropic/etc. official client libraries) into any crate. - MCP (Model Context Protocol) server calls are an explicit, scoped exception to the "no networked model calls" rule above. This carve-out permits MCP client behavior (stdio subprocess + SSE/HTTP transport, server lifecycle, tool discovery/invoke) in a new
plotforge-mcpcrate that depends only onplotforge-schema+ transport dependencies;plotforge-agentdepends onplotforge-mcp, never the reverse (preserves theschema ← mcp ← agent ← studioDAG). MCP server calls are local-only configuration: the user-global~/.plotforge/mcp.json(third member of the providers/skills/prompts family) never enters project source, contracts, traces, or export packages; credentials are referenced indirectly bycredential_env_varname only, never stored as values. MCP tool-call arguments and results are a new provider-response surface — they must pass throughcontains_secret_marker_text/redact_trace_textand enter traces only as redaction-safe summaries (tool name, server id, content-block kind, content hash), never raw bodies; anmcp_tool_call_hash(derived from non-secret config fields, mirroringprovider_config_hash) is carried in reproducibility metadata. Thetokioasync runtime is scoped toplotforge-mcponly (for SSE streaming);plotforge-agentand all other workspace crates must not pulltokioas a direct dependency —plotforge-mcpexposes a blocking façade (McpToolRegistry) to the agent. MCP tool-use integration uses a newMcpToolClientport trait + acomplete_with_mcp_toolswrapper orchestrator; the existingTextModelClienttrait and its three production impls are NOT extended or modified (the wrapper runs the multi-turn tool-call loop outsidecomplete_text_agent_output). Transport failures are explicit (mcp_spawn_failed,mcp_handshake_failed,mcp_tool_error,mcp_transport_unsupported,mcp_unknown_server,mcp_io); no silent fallback to a no-MCP path. Per-project MCP enablement persists via anenabled_mcp_servers: Vec<String>field onAgentSessionConfig(with#[serde(default)]for backward compatibility), mirroring theenabled_skillspattern. The CLI exposes MCP server management via a thinmcpsubcommand (orchestration only, never reimplements transport logic). MCP UI copy must not promise tool outcomes, automatic publishing, platform approval, or legal conclusions (no_launch_promise_lint.pyenforced). - Moderation provider calls are an explicit, scoped member of the real-provider family.
plotforge-agent::providers_moderationowns the independentModerationProviderport,OpenAiModerationClient, and blocking hand-writtenreqwestadapter; vendor SDKs and a directtokiodependency remain forbidden, andtokiostays scoped toplotforge-mcp.moderation_loop.rsperforms at most one optional pre-flight over the exact player input before Studio enters either its one-shot text or MCP-assisted text path; it does not widenTextModelClient, call text/MCP providers itself, or own a retry loop.local-piremains the only no-network text path and never loads or invokes moderation configuration. Moderation entries live only in the user-globalProviderRegistry.moderation_providersat~/.plotforge/providers.json;ModerationProviderEntry.credential_env_varstores an environment-variable name, never a credential value, and persisted entries are revalidated at the runtime build boundary before any network call. Successful response bodies are capped at one MiB, scanned withcontains_secret_marker_textbefore parsing, and normalized category labels are scanned again before exposure; traces/contracts receive only the schema-owned summary plus a redaction-safemoderation_config_hash, never a raw response, prompt, credential, or secret marker.JobKind::ModerationGeneration,UsageKind::Moderation, andReproducibilityMetadata.moderation_config_hashare additive contracts, but the pre-flight loop must not grow a parallel job state machine. RPM/concurrency use the shared provider throttle registry keyed by a canonical request endpoint identity so equivalent URL spellings cannot reset rate debt; moderationdaily_token_budgetis rejected while upstream responses lack trustworthy output-token accounting. When no moderation provider is configured, pre-flight is an explicit pass-through with no moderation hash; a configured provider failure or flagged result is explicit and never silently falls back to an unscreened text call.
Testing Policy
- Every new feature must include corresponding tests in the same change.
- Every bug fix should include a regression test that fails before the fix when feasible.
- Schema or serialization changes must include roundtrip/contract tests in
plotforge-schemaand affected integration tests. - Rule, runtime, storage, export, media, job, storycraft, and agent behavior changes must include crate-level tests for the changed boundary.
- CLI command changes must update black-box tests under
crates/plotforge-cli/tests/. - Export/player behavior changes must update export tests and, when rendering or interaction matters, the HTTP smoke path.
- Creator desktop changes must run
npm run creator-desktop:qa; user-visible UI adapters should add or update TypeScript/Vitest coverage for contract-backed behavior. - Creator desktop Vitest coverage is behavior-based: assert via role, accessible name, and visible text (
getByRole/getByLabelText/getByText), not DOM class names or internal state, so tests survive className/layout refactors and reflect what users actually perceive. - Static player changes must run
npm run player-web:qaand cover DOM interaction, mobile viewport behavior, and no-network package constraints. - Generated starter projects, traces, caches, and exports must stay out of commits unless they are intentional fixtures.
- CLI/export smoke tests must use temp dirs; do not mutate checked-in fixtures in CI.
- If a meaningful test cannot be added, document the reason in the PR or final response and run the next best validation.
Test Layer Boundaries
Tests live in four layers with non-overlapping responsibilities. Add a new assertion to the lowest layer that can express it; do not duplicate a business assertion across layers.
- Rust unit tests (
#[cfg(test)]modules insrc/, e.g.src/tests.rs): assert pure crate-internal logic only. No process spawns, no real filesystem writes outsidetempfile. This is the only layer that asserts the behavior of a rule, runtime transition, storage edit, or schema rule (e.g. "continue does not advance turn", "secret marker rejected"). - Rust integration tests (
crates/*/tests/*.rs): assert cross-crate collaboration through pub APIs. Usecreate_project_from_requestand other public storage/runtime APIs to build fixtures. Assert composition (e.g. "storage + export produces an audited package tree"), not business results already covered by unit tests. Provider HTTP integration tests live incrates/plotforge-agent/tests/provider_integration.rsand are gated with#[ignore]socargo testand CI skip them; run explicitly withcargo test -p plotforge-agent --test provider_integration -- --ignored. They exercise the full provider pipeline (text / image / TTS / moderation) against in-processstd::net::TcpListenermock servers (thecapturing_serverpattern), verifying request body wire shape, response parse, and error handling through the public client/adapter types — no real provider is contacted. APLOTFORGE_LIVE_TEST=1-gated stub is reserved for an opt-in real-API smoke and is always skipped in CI. - CLI smoke (
crates/plotforge-cli/tests/cli_smoke.rs): black-box tests of the compiled binary viaCARGO_BIN_EXE_plotforge-cli. Each test covers one command group and asserts "call succeeded + output shape" (exit code, presence of expected substrings, generated file existence). Do not re-assert business results that Rust unit tests cover. One mega-test that serially chains every CLI command is forbidden — split by command group so a failure in one group does not block the rest. - Frontend Vitest (
apps/creator-desktop/src/*.test.tsx,apps/player-web/test/*.test.js): assert rendering and user-perceived interaction (DOM, accessibility roles, visible text, click → state change).StudioDataSourcemocks must return contract-typed values (PlayOnceReport,RuntimeSnapshot, etc.) sonpm run typecheckcatches structural drift; never re-implement Rust business logic (trim/split rules, state-machine transitions, rule evaluation) inside a mock. To control mock behavior drift, rely on review discipline plus the Computer Use smoke — there is no low-cost automated guard.
Cross-layer dedup rule: the same business behavior (e.g. "continue keeps turn=0", "export excludes traces") is asserted in exactly one layer — the lowest layer that can express it. Upper layers may assert only that the call passed and that the returned value has the right shape, not that the business result is correct.
Required Validation
Use the narrowest relevant checks during development, then run the full gate before merging broad changes.
cargo fmt --all -- --checkcargo check --workspacecargo test --workspacecargo clippy --workspace --all-targets -- -D warningsnpm run creator-desktop:qanpm run player-web:qapython3 scripts/qa/creator_desktop_build_smoke.pyscripts/qa/full_local.shscripts/contracts/check_contracts.sh
Tempdir CLI smoke commands:
tmp="$(mktemp -d)"cargo run -p plotforge-cli -- new project --path "$tmp/starter-project" --force --concept "A local starter project." --visual-style "clear readable test style" --initial-scene "A creator opens a fresh PlotForge project."cargo run -p plotforge-cli -- check "$tmp/starter-project"cargo run -p plotforge-cli -- play "$tmp/starter-project" --oncecargo run -p plotforge-cli -- trace inspect "$tmp/starter-project/traces/latest.json"cargo run -p plotforge-cli -- export static "$tmp/starter-project" --out "$tmp/export"python3 scripts/qa/static_export_http_smoke.py --export-dir "$tmp/export"
Fixture validation commands:
- No committed default project fixture is expected.
CI and QA
- GitHub Actions (
.github/workflows/ci.yml, triggers onmainandmvp/**branches) keeps separate jobs forrust-static(fmt + no-launch-promise lint + check + contracts + clippy),rust-unit(needs rust-static),cli-smoke(needs rust-static), andexport-smoke(needs rust-static). - The
creator-desktopjob runs bothnpm run creator-desktop:qaandnpm run player-web:qa; there is no standaloneplayer-webCI job. scripts/qa/full_local.shis the repeatable local quality gate.- Creator Desktop build smoke lives at
scripts/qa/creator_desktop_build_smoke.pyand must stay wired intonpm run creator-desktop:qa. - Creator Desktop Browser/Computer Use smoke lives at
scripts/qa/computer_use_creator_desktop.md; verify the Studio shell, source editor, playtest run, runtime trace id, narrative review, diagnostics, and mobile viewport reachability. - Codex Desktop Computer Use smoke is local desktop QA only and must not become a GitHub Actions dependency.
- For Computer Use export smoke, serve static exports over localhost HTTP, then verify visible title, scene text, choice buttons, and post-click text changes.
- Starting a server is not enough evidence; perform at least one real browser/app interaction when Computer Use validation is requested.
Error Handling and Security
- Do not add silent fallbacks. Any fallback must be visible in trace/debug output and covered by tests.
- Prefer explicit errors over broad catch-all handling.
- Do not hardcode secrets, API keys, tokens, provider credentials, or personal paths into source, fixtures, tests, exports, or docs.
- Static exports must reject obvious secret markers and must not copy private traces or provider configuration.
- Static export package output must be audited against an explicit whitelist of player files, manifest files, and referenced assets; stale or unexpected files in the output package are errors.
- Do not introduce hidden network calls in tests or MVP runtime paths.
- Reference-library imports must store source metadata, authorization/rights metadata, short summaries, and structure notes only; do not store large raw copyrighted bodies in fixtures or project files.
Change Discipline
- Prefer targeted diffs that preserve crate boundaries.
- Remove obsolete logic when replacing behavior; do not layer duplicate sources of truth.
- Keep generated artifacts out of commits unless they are intentional fixtures.
- Update README and relevant docs when commands, workflow, or user-visible behavior changes.
- Keep
AGENTS.mdupdated when a rule becomes durable project memory.
