Imported from sandbox-social/silisocs (
AGENTS.md). Install upstream withnpx skills add sandbox-social/silisocs. Copyright stays with the author.
AGENTS.md
This file is a contributor guide for LLM coding agents working in this repository.
1) What This Repository Is
Silisocs is a native social simulation framework with an optional Concordia compatibility bridge for legacy scenarios. It has:
- YAML-first scenario and runtime configuration (Hydra + OmegaConf)
- A social-media game-master/environment layer
- Multiple platform backends (Twitter-like, Reddit-like, Mastodon)
- Declarative persona pipeline plus custom builder extension path
- Probe-based evaluation and rich runtime telemetry
- Studio visual workspace for scenario creation, launch, inspection, and analysis
The runtime entrypoint is:
src/silisocs/runtime/runner.py
2) High-Level Architecture
Core runtime layers:
1. Agent Construction Layer
src/silisocs/runtime/construction/agent_builders/- Builds agent construction specs from
agents.persona_pipelineand class data sources - Supports fixed-action set loading and template rendering
- Entry point:
AgentBuilder.build_agent_configs()
2. Agent Runtime Layer
src/silisocs/agents/base_agent.py— Abstract Agent interfacesrc/silisocs/agents/native.py— default LLM-backed native agentsrc/silisocs/agents/fixed.py— deterministic fixed-action agent- Custom agents subclass
Agent, accept aLanguageModel, implementname,observe(str), andact(ActionSpec) -> ActionOutput - To add a custom agent, point
persona_pipeline.classes.*.class_pathat the runtime class and provide strict constructorparams
3. Game Master Layer (Component-Slotted Architecture)
src/silisocs/environments/gm/base_game_master.py— Base coordinatorsrc/silisocs/environments/gm/game_master.py— ComponentGameMaster and MultiFlowGameMastersrc/silisocs/environments/gm/components/— Pluggable components:next_acting.py— Determine which agent acts nextobserve.py— Generate timeline/episode observationsresolve.py— Parse agent output into backend actionsapp_update.py— Schedule backend/recommendation updates
- To add custom component: implement
Componentinterface, set inenv.gm.components.{role}.class_path
4. Engine Layer (Execution Policies)
src/silisocs/simulation_engines/base_engines.py—RuntimeEngine, the single strategy-driven engine (startup, per-step GM updates, action concurrency, retry telemetry, probe phases). Scheduling/traversal is owned entirely by the step strategy;build_engineselects it fromsim.engine.step.built_in. Adding a new traversal means writing a step strategy and registering it — no new engine class.src/silisocs/simulation_engines/policies/— loop, step, and turn policies:- Turn policy:
single_action,fixed_count,open_ended - Step policy:
base,sequential,flow,multi_gm,multi_gm_serial,multi_gm_staged - Loop policy: default episode loop
- Turn policy:
- To add custom policy: implement the relevant policy ABC and reference it via
class_path
5. Backend Action Layer
src/silisocs/environments/backends/base.py— ActionCatalog, base app interfacesrc/silisocs/environments/backends/round_game.py— SimultaneousRoundGame, the reusable referee base for simultaneous-move repeated games (hidden choice buffering, resolve-at-the-round-boundary, payoffs, checkpoint round-trip)src/silisocs/environments/backends/twitter_like/— TwitterLikeApp with SQL backendsrc/silisocs/environments/backends/reddit_like/— RedditLikeAppsrc/silisocs/environments/backends/mastodon/— Real Mastodon server integrationsrc/silisocs/environments/backends/public_goods/— PublicGoodsApp, the reference game-theoretic backend (subclasses SimultaneousRoundGame)src/silisocs/environments/backends/messaging/— MessagingApp, the default agent-to-agent direct-message/broadcast channel (env=messaging)- Actions discovered via
@app_action(selectable_name=..., description=...) - To add custom backend: subclass
SocialBackendApp, implement action methods, register in app factory
6. Runtime Orchestration
src/silisocs/runtime/runner.py— CLI entrypoint, Hydra config compositionsrc/silisocs/runtime/configuration/— Config projection, external config loading, and validationsrc/silisocs/runtime/execution/session.py— Runtime assembly, simulation execution, checkpoint save/resume, and artifact writing- Handles: model creation, direct runtime construction, initialization, simulation execution, checkpoint save/resume
- Failure policy: prefer failing loudly. A condition the run cannot
legitimately continue past raises (config errors raise at build/first use, not
behind an invented default); what a run genuinely survives — one isolated agent
turn, one failed harness tool call, one routing call that fell back — is COUNTED
instead, never silently swallowed. Those counters are registered once in
evaluations/vocabulary.py::HEALTH_COUNTERS, which drives the run-end degraded warning,run_manifest.json'shealthblock, andRunArtifact.healthtogether (seedocs/usage.md→ "Run Health"). Add a new counter there, and it surfaces everywhere; emit it withSimMetricsCollector.get().increment_counter(...). effective_config.yaml(both copies) is written with everyapi_keymasked, so a run directory stays shareable; nothing reads credentials back from it.
3) Configuration Model
Top-level config composition (src/silisocs/conf/experiment.yaml):
- Defaults:
world: default,agents: default,sim: base,env: twitter_like,eval: base
Config groups and their base files:
| Group | Base file | Controls |
|---|---|---|
world |
world/default.yaml (@package _global_) |
Run params, setting, event, data |
agents |
agents/default.yaml (@package agents) |
Persona pipeline, shared memories |
sim |
sim/base.yaml (@package sim) |
LLM, engine, tool-calling, memory, checkpoint |
env |
env/twitter_like.yaml (@package env) |
Backend, GM components, initialization |
eval |
eval/base.yaml (@package eval) |
Probe configuration |
Key sim knobs (src/silisocs/conf/sim/base.yaml):
| Parameter | Value | Notes |
|---|---|---|
sim.llm.name |
gpt-4o-mini | Default LLM model |
sim.llm.temperature |
0.5 | Sampling temperature |
sim.llm.disabled |
false | No-op model for testing |
sim.action_mode |
custom | Prompt style (custom or generic) |
sim.tool_calling.mode |
single | none | single | multi |
sim.engine.step.built_in |
base | base, sequential, flow, multi_gm, multi_gm_serial, or multi_gm_staged (the three multi_gm* strategies select the flow-chain traversal mode; see below) |
sim.engine.turn_policy.built_in |
single_action | single_action | fixed_count | open_ended. fixed_count params: count; optional count_committed: true counts only actions that COMMITTED a backend change (a failed-resolve tool call no longer burns the budget), bounded by max_attempts (default 2*count) |
sim.engine.executor |
threads | Turn executor: threads (one pool worker per in-flight turn) | asyncio (turns are coroutines on one background event loop — thousands of LLM calls in flight on a handful of threads). Sync act/sample_text stay the REQUIRED contracts; act_async / sample_*_async / turn-policy run_async are OPTIONAL loop-native overrides (shipped: NativeAgent, FixedAgent, OpenAI-compatible providers, all built-in turn policies), and sync-only agents/models/policies automatically run on helper threads via asyncio.to_thread — both kinds mix freely in one step. max_concurrent_actions keeps its meaning (max in-flight turns, enforced as an asyncio semaphore); all scheduling semantics (flow chains, barriers, per-GM caps/locks, failure isolation, retry telemetry) are identical across executors |
sim.engine.control.built_in |
none | Interactive run control (play / pause / step one episode): none (default — no gate attached, loop runs uninterrupted, zero cost) | stdin (terminal REPL: n/next [k], c/continue, p/pause, s/stop between episodes) | control_file (obeys a JSON file another process writes: {"target": <int|null>, "stopped": <bool>} — target = first episode index to hold before, null = run freely). Implemented as a thread-safe StepGate (simulation_engines/control.py) the loop consults at each episode boundary, exactly like probe_runner/interventions; controllers only call the gate's primitives so it is input-agnostic. A custom LoopStrategy owns interactive control the same way it owns probe timing: call loops.await_step_permission(engine, step) at its episode boundary and break when it returns False — it returns True immediately when no gate is attached, so a replacement loop inherits play/pause/step/stop for free and costs nothing when control is off. start_paused holds before episode 0; control_file defaults to <output>/run.control; poll_interval (0.3s) is the file poll cadence. Control acts at episode boundaries and every paused loop still checkpoints per step, so a paused/stopped run is resume-stable. Studio exposes it as Step/Play/Pause/End-run on the live view (interactive launch → control_file overrides + POST /api/jobs/{id}/control) |
sim.engine.participation.built_in |
all | Sim-level per-step roster filter: all | activity_probability | activity_markov (or class_path to a custom ParticipationPolicy). Filters who is in the step's roster BEFORE scheduling, every GM's next_acting (effective acting = participation ∩ next_acting), AND per-step GM updates — so per-step backend work (e.g. recsys refresh) is O(active); an update component that needs the population declares requires_full_roster = True. Stateless/seed-derived, so replay- and resume-stable. The GM next_acting slot keeps only env-derived built-ins (all_agents, fixed_order); a config naming the moved built-ins there gets a migration error. Default all = everyone acts every step (deterministic, turn-based); activity gating is opt-in via activity_probability/activity_markov — rates are keyed by agent name or sim role, and an agent matching NO entry fails the run at its first step (naming the unmatched agents/roles) rather than falling back to an invented probability; the explicit ways out are per-role rates, a global active_probability, or built_in: all |
sim.engine.step.params.gm_turn_policies |
{} | Per-GM turn policy map (gm_name -> {built_in|class_path, params}); applies under any step mode |
sim.engine.step.params.gm_concurrency_caps |
{} | Per-GM concurrency caps (gm_name -> int); effective per-GM = min(cap, sim.max_concurrent_actions) via a per-GM semaphore; empty = global cap everywhere; applies under any step mode |
sim.checkpoint.every_n_steps |
null | Checkpoint frequency (run_study.py sets 1 by default) |
sim.checkpoint.save.built_in |
monolithic_json | On-disk checkpoint layout: monolithic_json (one JSON per step) | sharded (manifest + NDJSON object shards + raw SQLite sidecar files; params.objects_per_shard). Restore reads both transparently; class_path accepts a custom CheckpointSaveStrategy (mirrors the restore slot) |
sim.telemetry.record_active_agent_names |
false | Retain per-episode active-agent NAME lists in sim_metrics.json (O(active×steps)); counts are always recorded |
eval.probes.deployment.sample_k / sample_fraction |
null | Cap probe targets per due step AFTER the include/exclude filters (hash-ranked per (seed, step, agent) — replay/resume stable); unset probes every filtered agent |
eval.probes.probes.<name>.deployment |
(inherits global) | Optional per-probe deployment block overriding eval.probes.deployment per field (schedule/filters/sample_k/sample_fraction/at); unset fields fall back to global, validated with the same rules (per-probe errors name the probe). Probes sharing a resolved target set still batch into one questionnaire call. Per-probe sample_k draws an independent subset (ranking additionally keyed by probe name) |
eval.probes.deployment.at (+ per-probe) |
pre_step | Loop anchor a probe fires at: pre_step (before a step) | post_step (after a step) | run_end (once after the run — the terminal measurement; ignores start_step/every_n_steps). Each probe_events.jsonl row records its anchor. A custom LoopStrategy owns probe timing by calling loops.run_probe_phase(...) at the boundaries it wants |
agents.persona_pipeline.classes.<class>.model accepts a scalar model name OR a
full {name, temperature, provider, api_base, api_key, extra_kwargs, disabled}
block overriding sim.llm per-field (unset fields fall back to global;
extra_kwargs replaces, not deep-merges); models are deduped by effective config.
Key run params live in world/default.yaml (at config root via @package _global_):
| Parameter | Default | Notes |
|---|---|---|
num_agents |
10 | Number of agents |
num_steps |
5 | Simulation episodes |
scenario_name |
default | Used in output path |
seed |
1 | Random seed |
Note: scenario world/default.yaml files REPLACE the base world group (Hydra
searchpath shadowing), so every scenario must re-declare the universal run
params (num_steps, seed, output_rootname, ...). Flat scenario
sim.yaml/env.yaml files are MERGED into the composed config instead.
tests/test_bundled_scenarios_compose.py enforces this for the repo's example
scenarios. (Example scenarios live in scenarios/ and are repository content,
not packaged into the wheel; a pip install ships only the engine + base config.)
GM component routing is enabled with
env.gm.class_path=silisocs.environments.gm.game_master.MultiFlowGameMaster.
Scenario content lives under:
scenarios/<name>/conf/world/default.yaml(@package _global_) — run params + setting/event/datascenarios/<name>/conf/agents/default.yaml(@package agents) — persona pipeline- Optional:
scenarios/<name>/conf/sim.yaml— partial sim overrides (merged, not replaced) - Optional:
scenarios/<name>/conf/env.yaml— partial env overrides - Optional:
scenarios/<name>/conf/agents/thin.yaml— alternate agents variant (select withagents=thin)
For designing experiments via config (no code changes): See agent_docs/scenario_design.md
For understanding config structure deeply: See docs/configuration.md
For extending Studio analysis panels and views: See docs/analysis_panels.md
4) Defining New Agent Behaviors Cleanly
Use class-level behavior flows instead of adding custom manager branches:
- Assign class flow:
persona_pipeline.classes.<class>.flow_tag
- Define flow order:
sim.engine.step.params.flow_order
- Optional per-entity override:
sim.engine.step.params.agent_to_flow
- Optional observe specialization for selected flows:
env.gm.components.observe.params.episode_observation_flows
- Optional per-flow turn policy (how many actions a flow takes per step):
sim.engine.step.params.flow_turn_policies(map flow_tag ->{built_in|class_path, params}; unlisted flows usesim.engine.turn_policy). Only effective underengine.step.built_inofflow/multi_gm.- Optional per-GM turn policy (per-backend action cadence):
sim.engine.step.params.gm_turn_policies(map gm_name -> turn policy slot, same{built_in|class_path, params}shape; precedence per-flow > per-GM > global; resolved per batch by GM name, so it applies under ANY step mode). Lets a GM/backend set its own cadence — e.g.single_actionin a "world" GM butopen_endedin a social GM — and disambiguates multi-GM-chain hops that share one flow (whichflow_turn_policies, being flow-keyed, cannot). Unset/empty = the global turn policy applies everywhere (unchanged behavior). - Optional per-GM concurrency cap (how many of a GM's turns run AT ONCE):
sim.engine.step.params.gm_concurrency_caps(map gm_name -> int) caps that GM's concurrent agent turns via a per-GM semaphore; effective per-GM =min(cap, sim.max_concurrent_actions)(the global stays the overall ceiling and per-GM default). Orthogonal togm_turn_policies(turn policy = how MANY actions a turn takes; this = how many turns run AT ONCE). Empty = global cap everywhere; applies under ANY step mode. Use to throttle a rate-limited backend below the global limit while other GMs keep running concurrently.
- Optional per-flow action enforcement (which actions a flow may execute):
env.gm.components.resolve.params.flow_action_filters(map flow_tag ->{enabled_actions, excluded_actions}; further-restricts the backend filter, enforced at resolve time, never blocksFINISHED). Works on the default GM. A filter name matching no backend action raises at build for the catalog-bound resolvers (generic_action/tool_calling);parsed_actionkeeps literal matching with a warning (custom-mode verbs are world-defined).
- Optional per-flow probe/eval targeting:
eval.probes.deployment.include_flows/exclude_flows(deploy probes only to agents in the named flow(s); empty = all agents).
- Advanced multi-GM orchestration (optional):
env.gm_orchestration.gms(eachgms[*]— and the single defaultenv.gm— accepts optional per-GMaction_modeandtool_calling(a scalar mode string,none/single/multi, sibling to the scalaraction_mode) overrides; unset falls back to the globalsim.action_mode/sim.tool_calling.mode. The retiredtool_calling_modekey and thetool_calling: {mode: ...}block form each raise a migration error pointing at the scalar spelling. Resolve compatibility is validated PER-GM against each GM's effective mode: an effectivesingle/multiGM must usecomponents.resolve.built_in: tool_calling, anoneGM must not.)env.gm_orchestration.flow_bindings.flow_to_gms(maps a flow to a GM or a strictly-increasing-sequenceGM chain; the only supported flow binding key). A chain entry may benull(an empty slot) — meaningful only undermulti_gm_staged— meaning the flow IDLES at that stage and resumes at its next non-null hop. Trailing empty slots are trimmed (no effect); an all-empty chain is rejected ("cannot be empty"); the strictly-increasing-sequencerule applies to the real (non-null) GMs only. Undermulti_gm/multi_gm_serialnull hops are simply dropped. A chain entry may also be a branch node —{branch: {router, choices}}— that routes each of the flow's agents to one ofchoices(alternative GMs) at that one stage. Therouteris a{built_in|class_path, params}slot built by the engine (policies/routers.py+build_router) into a plain callable — there is no base class to subclass (the signature is formalized by the structuralRouterProtocol, positional-only parameters):route(agents, gms, ctx) -> {agent name: chosen gm name}, whereagentsis the flow's agent objects (callagent.act(...)freely),gmsis{gm name: game master}(one per choice, in config order; readgm.backendfreely), andctxisRouteInfo(flow, step, seed)for replay-stable decisions. Aclass_pathmay point at a class (built withparams) or a plain function (paramsbound as kwargs) — the "a custom function chooses" seam. The engine runs the router when the flow's chain reaches the branch stage — after the flow's earlier hops have drained, so the router sees live backend state; undermulti_gm_staged, after the prior stage's barrier — in all threemulti_gm*traversals. The router call runs UNLOCKED; the follow-up per-chosen-GM turn selection runs under that GM's lock; and the engine validates the returned assignment (every agent covered, every GM a real choice). Shared hops before/after the branch still run once on every agent (not split per choice). Built-ins:random(RandomChoiceRouter, weighted, deterministic per(seed, flow, step, agent)) andagent_choice(AgentChoiceRouter— asks each agent to pick a GM via a CHOICE probe; paramsprompttemplate +on_invalid: random|first|raise;on_invalidalso covers a routing call that RAISES — provider outage/retry exhaustion — so one agent's transient model failure only aborts the run underraise. Every fallback is TRACKED, not silent: it increments therouting_fallbacksrun-health counter, so it surfaces in the degraded-run warning,sim_metrics.json, and the manifest). An LLM-driven router is only as reproducible as the model it calls. Routing calls run serially on the chain-driver thread, outsidemax_concurrent_actions/gm_concurrency_caps— keep LLM-routed branches to modest flows. Under concurrentmulti_gm, branch hops into a shared GM with a STATEFULnext_acting(e.g.fixed_order) are not replay-stable (driver timing orders them); usemulti_gm_serial/multi_gm_stagedwhen exact replay matters there. Rules: ≤1 branch/chain, ≥2 distinct known choices whosesequences sit strictly between the branch's neighbours,multi_gm*mode only, not inside aflow_orderflow. Restore + per-GMowned_flowstreat a branch as any of its choices viacollapse_flow_chains/flow_chain_gm_names.- The multi-GM flow-chain traversal mode is selected by
sim.engine.step.built_in(multi_gm/multi_gm_serial/multi_gm_staged); the formersim.engine.step.params.chain_executionknob has been removed (a config that still sets it raises aValueErrorwith a migration hint):multi_gm(concurrent, default — unchanged behavior): runs flows as independent pipelines through their GM chains. Distinct flows advance concurrently and serialize only on shared-GM overlap (the per-GM lock), withflow_orderflows running first as a strict serial prefix.multi_gm_serial(legacy row-major): each flow runs its full GM chain to completion before the next flow starts, one batch at a time, in deterministic flow-order.multi_gm_staged(column-major with a global per-stage barrier):flow_orderflows run first as a serial prefix (same as concurrent); then every remaining flow advances ONE STAGE AT A TIME — all flows' stage-N hops run concurrently and stage N+1 does not begin until ALL of stage N's turns finish. Empty slots (nullchain entries) let flows with different chain shapes stay stage-aligned. Tradeoff: the barrier can leave the worker pool idle at stage tails (a fast flow waits for slow flows); usemulti_gmwhen you don't need stage alignment.
Default UX rule:
- Keep users on the default
ComponentGameMaster, with advanced Studio controls off. - Only expose flow tags and multi-GM controls behind advanced mode.
All per-flow keys above are additive and backward compatible: each falls back to
the global/default behavior when omitted, and the agent->flow mapping is the same
materialized agent_flow_tags used by scheduling and component routing.
Fixed agents (silisocs.agents.fixed.FixedAgent) are the reference example.
4.5) Mid-Run Interventions
An optional top-level interventions schedule (see
docs/configuration.md → "Mid-Run Interventions") fires actions at step
boundaries: set_participation, ban_agents/unban_agents,
set_component_params, set_recsys, set_turn_policy, swap_component
(persistent state-setters, replayed on resume for at_step < start_step),
inject_action, inject_post,
broadcast_observation (one-shot events, never replayed), or custom
(class_path to an InterventionHandler subclass). inject_action is the
generic injection language (any backend catalog action as a typed tool call
through the GM's resolve component); inject_post is sugar over it. Dispatch lives in
simulation_engines/interventions.py and runs at the single-threaded loop
boundary in FixedStepsLoopStrategy, so handlers mutate engine/GM/backend state
without locks. Fired-ness is a pure function of (schedule, step) — no
checkpoint-schema change. A ban is a BanFilterParticipation wrapper over the
active participation policy (soft ban), not a roster mutation.
set_component_params is the generic component-retuning language: a component
lists mid-run-safe parameter names in its class-level runtime_tunable
frozenset (BaseComponent.set_params routes each through set_<name>() or the
same-named attribute), and set_recsys is sugar over it. set_turn_policy
rebuilds a turn policy (global / per-gm / per-flow) into the maps the
scheduler reads each batch; set_router re-points a flow's branch router (a
stateless plain callable) in the step strategy's flow chain (multi_gm* only);
flow targets on both are preflight-validated against the flows statically
declared in config (deferred to fire time when none are declared);
swap_component hot-swaps a STATELESS observe /
next_acting / update component via the GM's rebuild_component seam (refused
if the outgoing or incoming component has non-empty get_state() — retune
stateful components with set_component_params instead; resolve/action_prompt
and initialize are out of scope). To add a kind,
subclass InterventionHandler (declare kind, persistent, validate,
apply) and reference it via kind: custom. Interventions fire BEFORE run_step
stamps the episode index, so a custom handler that logs backend events must stamp
first via the InterventionContext helpers: ctx.stamp_episode(gm) (stamps the
GM's action/exposure/harness loggers with ctx.step) or
ctx.resolve_action(gm, agent_name, output) (stamps, then resolves an injected
action through the GM's resolve pipeline — the seam the built-in inject_action/
inject_post handlers use).
5) Agent Interface: Concordia vs Custom
All agents in silisocs implement a common interface defined by silisocs.agents.base_agent.Agent (ABC).
Minimum Required Interface
Every agent (whether Concordia-based or custom) must implement:
from silisocs.agents.base_agent import Agent
class MyAgent(Agent):
@property
def name(self) -> str:
"""Return agent's display name."""
return "Alice"
def observe(self, observation: str) -> None:
"""Receive environment observation from the social app."""
self._last_observation = observation
def act(self, action_spec) -> str:
"""Generate an action response given the action specification."""
# action_spec provides context and constraints
# Return format is determined by the resolve component (via YAML config)
return "some action response"
The resolve component and agent's configuration determine the output format expected. Agents should not be concerned with prescribing action format—that is a platform concern.
Optional Async Path (sim.engine.executor: asyncio)
Sync act() is and remains the required contract. Additively, every agent MAY
override act_async(action_spec) (the base Agent default runs the sync act
on a helper thread via asyncio.to_thread), and every LanguageModel has
sample_text_async / sample_choice_async / sample_tool_calls_async /
sample_structured_async / sample_float_async twins (default: the sync method
on a helper thread; OpenAI-compatible providers implement text/choice/tool-calls
natively on an AsyncOpenAI client).
Under the asyncio executor each turn is a coroutine on one background event
loop: an agent with act_async acts loop-native (thousands in flight at once),
a sync-only agent's act runs on a bounded helper thread — blocking-safe, and
unable to stall the loop — and both kinds mix freely in the same step. To make a
custom agent loop-native, override act_async and route through the base
helper _call_model_async(context, action_spec) (same dispatch as
_call_model, awaiting the model's *_async twins):
class MyAgent(Agent):
def act(self, action_spec): # required, sync floor
return self._call_model(self._context(), action_spec)
async def act_async(self, action_spec): # optional fast path
return await self._call_model_async(self._context(), action_spec)
Everything that calls agent.act(...) synchronously (probes, seed-post
initialization, the agent_choice router) is unchanged. Custom TurnPolicy
classes may likewise provide an optional run_async (awaiting
engine.run_agent_step_async); a sync-only policy runs whole-turn on a helper
thread under the asyncio executor.
Reference Implementation: FixedAgent
silisocs.agents.fixed.FixedAgent is a concrete example of a non-LLM agent:
# src/silisocs/agents/fixed.py
class FixedAgent(Agent):
"""Deterministic agent executing pre-defined actions by episode."""
def observe(self, observation: str) -> None:
# Extract episode number from observation
self._current_episode = extract_episode(observation)
def act(self, action_spec) -> str:
# Look up action for current episode and return as string
action = self._next_action_item()
return format_action(action)
How Custom Agents Are Loaded
-
Create a native runtime class:
# my_agents.py from silisocs.agents.base_agent import Agent from silisocs.runtime.types import ActionOutput class MyCustomAgent(Agent): def __init__(self, *, model, name: str, context: str = ""): super().__init__(model) self._name = name self._context = context @property def name(self) -> str: return self._name def observe(self, observation: str) -> None: self._context += f"\n{observation}" def act(self, action_spec): return self._call_model(self._context, action_spec) -
Reference in world config:
persona_pipeline: classes: influencer: count: 1 class_path: my_agents.MyCustomAgent params: name: Alice context: Initial persona text. -
The runner imports the class path directly and instantiates it with
modelplus the configuredparams.
Concordia Integration Points
If building a Concordia-compatible agent (using EntityAgentWithLogging):
- Agents are context components (observe/act participate in component orchestration)
- Extend
EntityAgentWithLoggingto get logging + checkpoint support automatically - Opt in with
compat: concordia; the adapter calls the upstream Concordia prefab'sbuild()method
If building a custom (non-Concordia) agent:
- Implement only the
Agentinterface - Concordia integration still works (engine calls
agent.observe()andagent.act()) - Checkpoint support optional (implement
get_state()/set_state()if needed)
No special ABC requirement for Concordia agents—they naturally implement the interface via activity slots.
Tool-Calling Implementation for Entities
When tool-calling is enabled (tool_calling.mode: single|multi), the platform uses backend
actions as tools and the language model selects which action(s) to invoke.
Architecture for tool-calling:
- Detect tool-calling mode: The GM's action-prompt component checks the
enable_tool_callingflag and, when the backend providesgenerate_tool_schemas(), builds anActionSpecwithoutput_type=OutputType.TOOL_CALLSand the tool schemas inextra_args["tools"](seeenvironments/gm/components/action_prompt.py) - Agent act layer:
Agent._call_model()routesOutputType.TOOL_CALLSspecs to the model'ssample_tool_calls()with the provided schemas - Typed result: The agent returns an
ActionOutputcarrying typedToolCallentries (no string marker or JSON parsing involved) - Resolve execution:
ToolCallingResolveComponentvalidates the tool calls and executes them viabackend.invoke_action_with_kwargs()
Enabling Tool-Calling
To enable tool-calling at the game-master layer, configure resolve as tool_calling.
This keeps prompt generation mode (custom or generic) independent from parsing mode.
Example custom-cta + tool-calling:
sim:
action_mode: custom # custom prompt text still used
tool_calling:
mode: single
components:
game_master:
resolve:
built_in: tool_calling
When tool-calling is active, the native ActionSpec carries tool schemas in
extra_args; Agent._call_model() routes the request to sample_tool_calls.
Adding a Custom LLM Provider
Common providers that expose an OpenAI-compatible API ship as built-in presets
(anthropic, gemini, openrouter, groq, together, deepseek, mistral,
fireworks, xai, ollama); set sim.llm.provider to the name and supply the
key via the provider's env var (see OPENAI_COMPATIBLE_PRESETS in
runtime/language_models/factory.py). For anything else, two paths exist, with no
core edits required:
- Register a provider name (import the module before the run starts):
then setfrom silisocs.runtime.language_models.registry import register_llm_provider @register_llm_provider("my_provider") class MyModel(LanguageModel): ...sim.llm.provider: my_provider. - Or set
sim.llm.provider: mypkg.models.MyModel(fully qualified class path).
Providers that speak an OpenAI-compatible HTTP API should subclass
OpenAICompatibleLanguageModel to inherit retry/backoff and telemetry support.
Validation & Error Handling
Game master initialization (src/silisocs/environments/gm/game_master.py) validates agents:
# Checks at GM build time:
for entity in self.entities:
assert hasattr(entity, 'name'), f"{entity} missing 'name' attribute"
assert hasattr(entity, 'observe'), f"{entity} missing 'observe' method"
assert hasattr(entity, 'act'), f"{entity} missing 'act' method"
Runner validates direct class construction, so missing methods fail fast.
Multi-Action Support (Open-Ended Policy)
When using sim.engine.turn_policy.built_in: open_ended:
- Agent's
act()method is called repeatedly within one step - Agent should output valid actions OR the special "Finished action episode" signal
- Resolve components recognize "FINISHED" and stop iteration
- Allows agents to decide how many actions to take per step
Example:
def act(self, action_spec) -> str:
if self._done_for_this_step():
return "Finished action episode"
return self._next_action()
This mode works with any agent (Concordia or custom) that implements the basic interface.
Harness Agents (real agent harnesses as silisocs agents) — EXPERIMENTAL
Experimental. The harness integration is new and evolving; the deterministic core (Tool Bridge, agent/probe/checkpoint contract, Model Proxy) is tested with the dependency-free
FakeHarnessAgent, but live Hermes/OpenClaw runs are opt-in and less exercised. Expect the config surface and internals to change.
A harness agent (silisocs.agents.harness) embeds a real agent harness — Hermes
(in-process) or OpenClaw (out-of-process) — as an Agent. The harness runs its own
model→tool loop inside ONE act() call (bounded by its max_iterations), so for a
harness class the effective turn policy is single_action around one complete run. The
design is deliberately thin at the edges:
- Harness agents are just
Agentsubclasses loaded viaclass_path. No engine, scheduler, participation, intervention, or checkpoint changes — those layers are duck-typed.HarnessAgentowns observe-buffering, ActionSpec dispatch, probes, checkpoint state, and telemetry. - One thin module — the Tool Bridge (
ToolSurface). Built per turn from the acting GM's backend. It adds no filtering of its own — the backend already owns that:schemas()returnsbackend.generate_tool_schemas()(already restricted to the backend's enabled/excluded actions) andexecute()forwards tobackend.invoke_action_with_kwargs()(which validates the call and logs theaction_events.jsonlrow, just like a native turn). The surface only adds the harness concerns: injecting the actor argument (viabackend.action_accepts_param, shared with the resolve components), turning exceptions into results the loop reacts to, and recordingharness_events.jsonl. It is the single seam every adapter consumes, so harness agents are backend-agnostic by construction. - One adapter seam —
HarnessAdapter(Protocol). Concrete harnesses implement onlyrun_turn(optionallyrun_turn_async/run_probe/snapshot/restore/bind_model_proxy).FakeHarnessAdapter/FakeHarnessAgentare the deterministic, dependency-free reference (and the contract-test subject). - Zero GM config (self-describing). No
harnessGM built-ins exist. The default action-prompt component binds the Tool Bridge intoActionSpec.extra_args['tool_surface']for any acting agent that declareswants_tool_surface(harness agents do; native agents don't, so non-harness runs are unchanged). A harness turn'sActionOutputcarries aharness_turnstructured payload, so the shared_BaseResolveComponent.resolve_actionrecords it uniformly — ANY resolve (parsed_action/tool_calling/…) handles harness output, and one GM hosts mixed native+harness populations with no special config or validation. - One model plane — the Model Proxy (
agents/harness/proxy.py). A loopback OpenAI-compatible server the runtime starts when any harness class is configured (setup_harness_proxyinagents/harness/runtime.py, stopped in the sessionfinally). Harness model calls forward through it byte-for-byte; providerusagefolds into the SAMEllm_usageas native agents via a duck-typedUsageAccumulatoradded to the run'smodels. Real keys live only in the proxy; per-agent routing tokens are the harness-side "api key". - Determinism: harness agents are snapshot-restored, never replay-restored. Per-call
detail is written to
harness_events.jsonl(per-GM in multi-GM runs), indexed in the run manifest and exposed onRunArtifact.iter_harness_events().
To add a new harness: implement a HarnessAdapter and a thin HarnessAgent subclass.
See docs/harness_agents.md. Non-goals of the current pass: no event-driven engine, no
MoltBook facade, no live-internet tools (surfaces expose only backend catalogs).
5.5) Action Modes and Platform Configuration
The platform supports different action modes configured via sim.action_mode:
sim:
action_mode: custom # custom | generic
tool_calling:
mode: none # none | single | multi
Each mode corresponds to how the agent's responses are interpreted and executed:
- custom: Custom parsing format determined by the world
- generic: Generic action name + parameters format
Tool-calling is configured separately via sim.tool_calling.mode.
The specific action format and response interpretation is determined by the resolve component and world configuration, not by the agent. Agents simply return strings; the platform interprets them according to the active mode.
For tool-calling mode specifically: The entity layer is responsible for calling sample_tool_call() when the action_spec indicates tool-calling is needed. The resolve component then processes the result. This architecture keeps tool-calling logic in the entity/act layer, not in resolve.
6) Checkpoints and Replay
- Checkpoints are saved as JSON under run output
checkpoints/step_{N}_checkpoint.json. - Runtime resume uses:
sim.checkpoint.source_run(explicit prior output directory)sim.checkpoint.auto_resume(defaulttrue): whensource_runis unset, resume from this run's own output directory if it already contains checkpoints; a fresh output directory still starts from scratch. Setfalseto force a fresh start.sim.checkpoint.restore- Resume restores game-master and entity component state plus raw log.
Custom backend contract (the enforced parts; full checklist in
docs/backends.md): implement name()/description() (abstract); expose actions
with @app_action and name the actor param agent_name for injection; return
ActionResult(committed=False) for rejected or idempotent calls (committed
calls are logged automatically); accept the factory's ctor kwargs or not
— action_logger is wired post-construction either way, so it can no longer be
silently swallowed. Validations that now fail at build rather than mid-run: an
action filter leaving no callable action (enabled_actions: [] is an empty
ALLOW-LIST, not "no filter"; FINISHED doesn't count); provides_checkpoint_state=True
with an inherited no-op get_state/set_state; and a social-only component
(social_media/timeline_every_turn/social_recommendation, or any component
declaring requires_social_backend = True) on a non-SocialBackendApp backend.
An omitted observe slot picks timeline_every_turn for a social backend and
app_observation otherwise. Backends need NO database (db_path belongs to the
SQLite backends; resource_market/virtual_space are in-memory references).
Open tags and semantic fields declared on @app_action are derived for
custom backend types and persisted in the run manifest. A class-level
event_semantics declaration remains available for aggregate definitions (an
EventSemantics or the portable {roles, fields, labels} mapping — the shipped
social backends declare theirs this way); resolution MERGES the declaration with
decorator-derived tags (declaration wins per entry), so decorating a new action
always reaches analysis, and a malformed class declaration raises.
Committed-only action log (backend authors): action_events.jsonl is the
canonical log of actions that committed a state change or performed a deliberate
logged read. The invocation layer logs plain returns and
ActionResult(committed=True) automatically; return
ActionResult(committed=False) for rejected or idempotent calls. log=False
disables logging for a deliberate read, log_as selects a stable label, and
data supplies derived logged fields. _log_action_event remains for
non-action commit points and direct dispatchers; calling it during an invoked
action suppresses the automatic row. invoke_action_detailed(name, kwargs) -> (committed, result) reports the outcome to committed-counting policies.
Every committed row also appends to an in-memory committed-events mirror
on BackendApp — the supported runtime read path so scenario code (routers,
intervention conditions, state-dependent policies) can ask "how many X committed
so far" without scraping the log file: count_committed_events(labels=..., agent=..., since_episode=..., before_episode=..., text_contains_any=...) and the
iter_committed_events(...) generator behind it (per-backend / per-GM; snapshot
backends round-trip it via the _committed_events_state()/_restore_committed_events()
helpers, replay-restored ones rebuild it for free).
Backend restore contract (backend authors): a backend supports either (or both)
of two restore paths. (1) Authoritative snapshot — set the class flag
provides_checkpoint_state = True and make get_state/set_state round-trip the full
state; restore applies it directly. (2) Action-event replay — this is a mechanism
that lives entirely in the pluggable sim.checkpoint.restore strategy, not on the
backend. The built-in social_action_event_replay strategy owns its per-backend
event→action mappings in a registry keyed by backend_type
(runtime/checkpointing/replay_mappers.py): a backend "supports replay" exactly
when a mapper is registered for its backend_type. Backends themselves carry no
replay method — they implement only get_state/set_state. The shipped registry
maps twitter_like to the stateless microblog_event_to_replay_action; reddit_like
has none (no valid microblog mapping) and self-restores via snapshot. A custom
backend opts into the built-in strategy with
register_replay_mapper(backend_type, mapper) (mirrors register_llm_provider) — no
core edit, no backend method; a bespoke restore that a stateless mapper can't express
is a custom sim.checkpoint.restore.class_path strategy instead. Every shipped
backend self-restores via set_state (provides_checkpoint_state=True); Mastodon —
which can't snapshot its live server — does so by embedding its action history in
get_state and replaying it (with old→new toot-id remapping) in set_state through
its own private mapper, so it is not driven by the replay strategy at all. The
authoritative-vs-replay decision is per game master: snapshot GMs restore from
their block, the rest go to the strategy (which requires a registered mapper).
Multi-GM runs isolate each GM's backend db + action_events.jsonl under
<output>/<gm_name>/; both restore and eval discover these via
silisocs.evaluations.action_events.resolve_action_event_files (restore passes
all per-GM logs to the strategy as action_events_files). A GM may override the
global sim.checkpoint.restore with its own strategy via
env.gm_orchestration.gms[*].restore (same schema; absent → global default).
Save layout (sim.checkpoint.save, mirroring the restore slot): monolithic_json
(default — one JSON per step, the long-standing format) or sharded (a
step_N_checkpoint.json manifest plus NDJSON object shards and raw SQLite sidecar
.db files verified by sha256; params.objects_per_shard, default 500). The payload
always comes from make_checkpoint_data; only the on-disk layout differs, and
load_checkpoint_file reassembles a sharded checkpoint transparently, so restore
strategies and eval.py never see the difference. A custom layout is a
sim.checkpoint.save.class_path subclassing CheckpointSaveStrategy
(runtime/checkpointing/save.py).
Saving policy: checkpoint saving is disabled by default when running directly
with the silisocs CLI unless every_n_steps or explicit_steps is configured.
When running via silisocs-study, checkpointing is enabled automatically
(every_n_steps=1) so that evaluators can access the final checkpoint for
action-type metrics. Studies can change the frequency via
run_defaults.overrides: {sim.checkpoint.every_n_steps: N}.
For custom agents: Checkpointing is duck-typed — every agent and game master is
checkpointed by reading its get_state() and applying its set_state() (there is no
isinstance(EntityWithComponents) gate). The base Agent provides no-op defaults, so an
agent with no episodic state needs no changes. If your custom agent does have episodic
state, implement BOTH get_state() and set_state(): restore now refuses to load a
checkpoint whose object saved non-empty state but only inherits the no-op set_state()
(it would otherwise be silently dropped). Restore also rejects a checkpoint object whose
class_path/compat no longer matches the current runtime object of the same name.
Example:
class MyAgent(Agent):
def get_state(self) -> dict[str, Any]:
return {"episode": self._current_episode, ...}
def set_state(self, state: dict[str, Any]) -> None:
if state:
self._current_episode = state.get("episode", 0)
7) Key Development Commands
Use uv-managed workflows (from docs/contributing.md):
- Sync dev env:
uv sync --group dev - Lint workflow:
uv run --group dev poe lint - Test workflow:
uv run --group dev poe test - Docs workflow:
uv run --group dev poe docs - Pre-commit hooks all files:
uv run pre-commit run --all-files --verbose - Commit with Commitizen:
uv run cz c
Fast contributor workflow (coding-agent friendly):
uv sync --group dev- Run targeted tests for changed files first (
uv run pytest <targeted_tests>) - Run full quality gate:
uv run pre-commit run --all-files --verbose - Run coverage workflow:
uv run --group dev poe test - Commit with Commitizen (
uv run cz c) or a validcz_gitmojimessage - Push branch (
git push origin <branch>)
NEVER commit .env files or any file containing API keys, passwords, or secrets.
These files (.env, store.env*, etc.) are gitignored for this reason. Staging them
accidentally (e.g. via git add -A) and pushing will expose credentials publicly and
trigger GitHub push protection. If you suspect a secret was staged, run
git reset HEAD <file> before committing.
8) Testing Expectations for Agents
When changing runtime behavior:
- Run targeted tests for touched modules first.
- Run full suite before finalizing if feasible.
- Add tests for new config/behavior paths.
- Avoid deleting tests unless they are obsolete due to architecture removal.
Useful tests in this repo include action parsing, worker limits, probe deployment, backend action catalogs, and checkpoint policy tests.
9) Documentation Map
This guide (AGENTS.md) is for any coding agent or human contributor who is
extending the framework — writing new components, backends, agents, or
changing architecture. CLAUDE.md intentionally points here so Claude Code,
Codex, Cursor, and other repo-aware agents share one canonical map.
If instead you want to design and run experiments via config only: → See agent_docs/scenario_design.md — Scenario design guide for config-based users
Agent-doc index: → See agent_docs/README.md — discoverable map of agent-facing guides and guided workflows
Detailed architecture deep dive (multi-flow, multi-GM, component routing): → See agent_docs/architecture.md — Reference for complex orchestration patterns
Guided workflows (interactive design workflows — readable by any coding agent): → agent_docs/skills/new-scenario.md — Step-by-step scenario design assistant → agent_docs/skills/new-study.md — Step-by-step study design assistant
Public documentation (for end users):
docs/index.md— Hub for all documentationdocs/configuration.md— Config reference (all knobs explained)docs/usage.md— End-to-end workflowdocs/study_schema.md— Study YAML schema, directory layout, and analysis conventionsdocs/environment_layer.md— Engine/GM/component extensibility patternsdocs/backends.md— Backend plugin patternsdocs/building_agents.md— Agent builder patternsdocs/studio.md— Studio usagedocs/contributing.md— Code standards
When adding features, update docs in:
- Config schema and fields (docs/configuration.md)
- Runtime behavior and extension guidance (this file + docs/environment_layer.md)
- User-facing usage examples (docs/usage.md)
- Studio behavior (
docs/studio.mdif applicable)
10) Common Pitfalls
- Adding GM/engine bloat instead of using flow routing + component hooks
- Breaking the action text format consumed by resolve
- Forgetting to keep docs aligned with runtime defaults
- Assuming Studio run-artifact loading equals checkpoint state replay
- Relying on non-uv environment when reproducing tests
- Not understanding config composition (Hydra merges scenario-local overrides with base defaults)
11) PR Readiness Checklist
- Code compiles and tests pass in uv environment
- Lint/pre-commit workflow passes
- New behavior has tests
- Docs updated for config + usage + architecture
- Commit message uses the configured
cz_gitmojischema (see §14)
11.5) Branching Rules
- Never commit code changes directly to
main. All code changes must go through adevbranch (or feature branch) and be merged via PR. - Documentation-only changes (edits to
docs/,agent_docs/,AGENTS.md,README.md) may be committed directly tomain. - When starting new work, create a branch:
git checkout -b dev(or a descriptive feature branch name).
12) Entry Points for Quick Exploration
Start from these files to understand the flow:
- Config composition:
src/silisocs/runtime/runner.py— How Hydra merges configs - Simulation orchestration:
src/silisocs/runtime/execution/session.py— Full workflow - Engine execution:
src/silisocs/simulation_engines/base_engines.py— Episode loop - Game master:
src/silisocs/environments/gm/game_master.py— Simple preset - Multi-flow GM:
src/silisocs/environments/gm/game_master.py— Advanced component routing - Component slots:
src/silisocs/environments/gm/components/— Pluggable behavior - Backend actions:
src/silisocs/environments/backends/twitter_like/app.py— Example backend - Run loading:
src/silisocs/evaluations/run_artifact.py—load_run/load_studytyped artifact loaders (manifest-first with legacy fallback); Studio and analysis tools load runs through this, never by rediscovering the file layout
13) Session State
Use a SESSION_STATE.md file (gitignored) to maintain context across a work session.
- At session start: check if
SESSION_STATE.mdexists and read it to restore context. - After significant subtasks (commits, refactors, feature completion): offer to update it.
- Clear when starting unrelated work or after a clean commit.
Template:
# Session State
## Current Focus
Brief description of current task
## Modified Files
- path/to/file.py - what changed
## Decisions Made
- Chose approach X because Y
## Next Steps
- [ ] Pending task 1
- [ ] Pending task 2
## Open Questions
- Question for user about X?
14) Environment Notes
- Use
uv runprefix for all commands. - Run pre-commit before committing (see section 7 for workflow).
- Commit messages must use the configured
cz_gitmojischema:♻️ refactor(...):— renames, restructuring🐛 fix(...):— bug fixes✨ feat(...):— new features📝 docs(...):— documentation🧹 chore(...):— maintenance
- WSL users: if imports are slow (1+ min), the venv is likely on
/mnt/c. Use a WSL-native venv:export UV_PROJECT_ENVIRONMENT=~/venvs/simulator uv sync
