Imported from zh30/xiaomaolv (
AGENTS.md). Install upstream withnpx skills add zh30/xiaomaolv. Copyright stays with the author.
AGENTS.md
Build & Test
# Format
cargo fmt --all
# Lint
cargo clippy --all-targets -- -D warnings
# Test
cargo test --all-targets
# Run integration test
cargo test --test <test_name> -- --nocapture
# Loop Engineering focused regression tests
cargo test --test harness_loop_engine -- --nocapture
cargo test --test http_loop_engine_api -- --nocapture
cargo test --test service_harness_trajectory -- --nocapture
cargo test --test harness_evolution_engine -- --nocapture
# Production build (thin LTO, opt-level 3, stripped)
cargo build --release
Architecture
xiaomaolv is an AI gateway service that routes messages through AI providers, manages conversation memory, and supports extensible plugin architectures.
Core Data Flow
Telegram/HTTP → Channel → Service → Memory → Provider (AI) → StreamSink → Channel → Response
↓
Scheduler (cron jobs)
↓
MCP Runtime / Skills / Code Mode
↓
Trajectories/Feedback → Evolution Engine → Shadow Eval → Human Promotion
HTTP/Telegram/Signals → LoopEngine → Workflow DAG → Leased Worker → Checkpoints/Artifacts
↓ ↓
Goal events/SSE Self-test / Replay
Key Traits
ChatProvider::complete()/complete_stream()- AI inference entry pointsStreamSink::on_delta()- streaming response handler (implemented by channel)ChannelFactory::create_channel()- channel instance creationProviderFactory/ChannelFactory- plugin registration pointsLoopStore- durable Goal/Workflow/Attempt/Checkpoint, Signal, Replay, Self-test, and Artifact persistence seamWorkHandler- registered Dynamic Workflow execution boundary with an explicit effect class
Message Processing Pipeline (service.rs)
The Service struct orchestrates the entire pipeline:
- Receives inbound message from channel
- Queries memory for relevant context
- Optionally runs MCP auto tool loop
- Optionally runs Skills matching
- Sends to AI provider
- Streams response back through channel's StreamSink
Memory Backend
- sqlite-only (default):
memory.rshandles all storage - hybrid-sqlite-zvec: Layered search combining SQLite + vector sidecar for semantic search
- Keyword fallback when vector search returns nothing
hybrid_min_scorethreshold gates relevance injectioncontext_memory_budget_ratiolimits memory as % of input budget
MCP Auto Tool Loop
When agent.mcp_enabled = true:
- Discover tools from configured MCP servers
- Ask model to emit JSON tool calls
- Execute tools, feed results back to model
- Stop at
mcp_max_iterationsor final answer
Skills Runtime
- Skills are local scripts loaded dynamically with semantic matching
skills_match_min_scoregates selectionskills_max_selectedlimits how many skills are injected into prompt- MCP tools and Skills operate at different levels: MCP extends AI capabilities, Skills extends agent behavior
Telegram Group Behavior
- smart mode (default): contextual auto-trigger with
group_followup_window_secsrecent bot context window; learns summon aliases automatically - strict mode: requires explicit @mention or reply
- Scheduler commands: natural language parsing with confirmation workflow
Code Mode
Safe-by-default execution layer with two modes:
local: direct Rust evaluation (limited to math, string, format ops)subprocess: spawns subprocess with resource limits
Capability metadata filters MCP access before execution. subprocess is not an OS-level sandbox.
Self-Evolving Harness
EvolutionEngineowns the discover/propose/evaluate/approve/activate/rollback state machine- Evolution is disabled by default and only evolves a bounded replacement system-prompt patch
- Automatic cycles consume failed trajectories or negative feedback and stop at
ready - Shadow evaluation calls the provider directly and cannot execute tools, write memory, or send messages
- Shadow evaluation scores operator eval cases plus the versioned benchmark suite (
harness/benchmark.rs,id@version); benchmark regressions are always fatal regardless ofmax_regressions - Human approval and activation are separate authenticated operations; rollback restores the prior deployment
- SQLite stores candidates, eval snapshots, feedback, deployments, the active pointer, and immutable audit events
- Full bounded evidence SHA-256 is globally unique to prevent stale or concurrent duplicate proposals
Loop Engineering Harness
LoopEngineowns the durableGoal -> Workflow revision -> WorkItem DAG -> Attempt -> Checkpointlifecycle- Planning never dispatches work; approval binds the exact goal revision, plan hash, effect manifest, acceptance criteria, and execution budget
LoopWorkeruses at-least-once claims, expiring leases, monotonically increasing fencing tokens, and prepared/committed/reconciled checkpoints/resumeexpires stale leases and reconciles committed outcomes without replaying their effects- Registered handlers are
goal_planner,provider_analysis,self_test_suite,session_replay,manual_gate,evolution_evaluate, and — only behindexternal_write_enabledplus a handler allowlist —channel_send; unlistedexternal_writehandlers are rejected at plan/approve/register/dispatch;waiting_confirmationparks are resolved via the goal-scopedresolve-confirmationroute (confirmed/retry/abandoned+ audit reason) - Multi-source
EvolutionSignalrecords preserve source/trust/deduplication metadata; ingestion dedups on exact keys, normalized content, and token-set near-duplicates; only an operator route can convertobserved/triagedsignals intoproposedgoals or mark themignored(HTTP + Telegram/signals) - Production
coreself-tests are read-only; repeated identical failure sets produce one deduplicated internal signal - Provider frames are recorded for main plain, Code Mode, and MCP completion paths when trajectory logging is enabled; structural replay executes zero live tools
- HTTP collection/detail resources plus per-goal monotonic SSE are the Desktop boundary;
GET /consoleserves an embedded single-file operator dashboard over that contract (same bearer auth, no new backend semantics) - The evolution adapter may evaluate an existing prompt candidate, but approval, activation, and rollback remain exclusively in
EvolutionEngine
Plugin Architecture
Provider (provider.rs + provider_plugin_api.rs):
- Implement
ChatProvidertrait - Register
ProviderFactoryinProviderRegistry - Built-in:
OpenAiCompatibleProviderFactory
Channel (channel.rs + channel_plugin_api.rs):
- Implement channel types in
channel/subdirectory (group_pipeline, update_pipeline, workers) - Register
ChannelFactoryinChannelRegistry
Configuration
TOML with env placeholders (${VAR:-default}). Key sections:
[app]- bind, default provider, locale, concurrency limits[providers.<name>]- AI provider configs[channels.telegram]- streaming, scheduler, group behavior[channels.http]- HTTP channel for programmatic messaging[memory]- backend selection and hybrid settings[agent]- MCP, skills, code mode settings[agent.harness]- trajectory, compaction, verification, evolution, and loop-engine settings[agent.harness.evolution]- self-evolution cycle, evidence, eval gates, and human promotion policy[agent.harness.loop_engine]- durable control plane, scoped signal ingestion, worker leases/concurrency, and maintenance interval
Current operator documentation: docs/loop-engineering-harness.md. Historical documents under
docs/plans/ describe decisions at the time they were written and are not the runtime contract.
