The engine's internals, its contracts, and where to extend it. Companions:
DESIGN.md (why it's shaped this way),
authoring.md (writing scenarios), and the test suite —
every contract described here is pinned by a test in tests/.
warren.core world · components · params DSL · effects · adjudication
warren.runtime sim clock · wake policies · durative plans · messages · processes
warren.space grid/room compilers · pathfinding · move / move_to
warren.perception visibility · observations · versioned text renderer · themes
warren.mechanics packs (inventory, resources, comms, damage, vehicles, production, encounters)
warren.author scenario micro-DSL + build-time validation
warren.agents runners (scripted / random / policy)
warren.viz trace recorder · PixiJS viewer
warren.protocol warren-agent/0 wire models (no engine imports)
warren.remote agent mailboxes · RemoteRunner · briefings/manifest builders
warren.serve FastAPI host (spectators + /v1 agent control plane)
warren.client httpx-only protocol client + framework adapters (pydantic-ai)
warren.analysis scenario-agnostic trace summaries
Dependencies point strictly downward. The kernel knows nothing about time policy, perception, or geometry; mechanics packs are ordinary scenario-level code that happens to ship in the library.
The heap is the only clock. When its head is a system event (durative
completion, plan step, message arrival, process tick, goal deadline) the
runtime drains everything due at that instant straight into adjudication —
no agent is consulted. When the head is a wake, it closes a batch that
never extends past the next system event, builds every observation against
the same world, gathers decisions concurrently, and applies them in agent
order. Both paths go through the same deterministic engine.step, whose
events feed the wake policy (who wakes next), goal evaluation, and the
trace. A fuller, per-decision-branch version of this figure lives in
media/lifecycle.mmd (SVG).
Entity(id, kind, template, name, desc, tags, components, state) — the only
mutable field is state. components holds frozen dataclasses registered via
@component("name"); they serialize by registry name, so import the module
that defines a component before deserializing worlds that use it.
Relations (src, kind, dst, plus props/state) carry all structure:
at (placement), holds (inventory), route (topology, with
state.open/state.blocked_by), and any scenario-defined kinds. The world
maintains indexes — by kind, by component, occupants-by-container, relation
adjacency both ways — so perception and affordances never scan all entities.
place_of() follows the containment chain (held item → holder → vehicle →
cell); container_of() is one hop. Mutation happens only through
World.apply(delta).
The design's core move. A param spec answers three questions from one declaration:
params={"item": EntityParam(comp="carryable", scope="colocated")}- validate a submitted arg (wrong kind/scope/filter → teaching rejection
like
item:not_here:gem), - enumerate the candidate domain for observations (from indexes, capped),
- schema for LLM tool calling.
ActionSchema.tool_view() builds the agent-facing tool: per-param choices
for finite domains, plus fully-validated concrete suggestions when the
combination space is small (≤200 combos, capped at 10 suggestions). A required
finite param with an empty domain hides the tool. Free-form params
(IntParam, StrParam) surface as schema only.
When adding a param type, implement domain / validate / schema and keep the
invariant: everything domain returns must pass validate.
Declarative effects (Set, Inc, Take, PlaceAt, Transfer, Spawn, …)
compute a Delta and report the loci they write. Custom Fx(fn, at=...)
must declare loci or it conservatively locks its actor. Loci shapes:
("state", id, field), ("place", id), ("own", id),
("rel", src, kind, dst, field), ("id", id).
Multi-effect actions resolve against a scratch clone so a failing effect can't
half-apply (see ActionSchema.resolve — swap in a COW overlay if world clones
ever show up in profiles).
Engine.step(actions, time):
- reject duplicates / unknowns; validate params
- compute loci; union-find proposals whose loci overlap into groups
- per group: schema-owned
GroupResolvers adjudicate jointly (e.g.Contest); everything else re-validates sequentially in actor order against the updated world — which is why a depleted well or a taken item rejects later contenders with an honest reason, and a capacity-2 train admits exactly 2 of 3 simultaneous boarders - apply deltas, append
Events to the log
Multi-effect actions resolve to a canonical net delta (relations as a
presence diff, mixed set/inc collapsed to final values): the authored effect
sequence is honored and replay matches live execution. World.apply preflights
every reference — a bad delta raises DeltaError before any mutation, which
the engine converts into a rejected event.
Determinism contract: world transitions are a pure function of
(genesis, ordered action stream, seed). Every RNG derives from
derive_rng(seed, *key). Scheduling policies change which actions get
proposed — histories differ across policies by design; tests pin the former,
never the latter. Corollaries: never call random/hash() in mechanics or
policies (use ctx.rng / zlib.crc32), and never mint ids from global
counters (see the battle-id regression test).
One heap queue in simulated ms. System entries (plan steps, deliveries, process ticks) sort before wakes at the same instant, and a wake batch never extends past the next system event (conservative causality: everyone in a batch decides on the same world).
The loop per batch: collect wakes in the window → build all observations against the same state → gather decisions (concurrently — async runners overlap within a batch) → apply in deterministic agent order. Caveat: in busy worlds batches are near-singletons (measured mean 1.54 agents on the 100-agent city run), so slow runners still serialize across batches. The planned fix — treating LLM thinking time as a simulated duration the scheduler can look ahead through — is on the roadmap.
Decisions: Act(name, args) | Sleep(ms) | None. What happens next:
| Case | Behavior |
|---|---|
| instantaneous action | batched into one engine.step at batch time |
duration= action |
validated now for fast feedback; adjudicated again at completion via engine.step (so completion-time contention is correct) |
plan= action |
validated now; steps execute on the clock, each re-checking its guard; failure emits aborted and frees the lane |
| busy lane | rejection + re-wake when the lane frees (lane_free), not a retry spin |
| wake while busy | observation carries busy (lane, action, args, until_ms); tools excludes busy-lane actions, so the agent is never offered an action it can only get rejected for |
| rejection | policy re-asks after rejection_ms (default 400ms), not 1ms |
Sleep(ms) |
swallows all wakes except messages until the deadline; each completion wake is stamped with its deadline, so re-sleeping after a message discards the superseded wake (no accumulation) while the current one still fires even when batching pulls it in early |
| simultaneous completions | default-durative actions finishing at the same instant adjudicate as one Engine.step cohort, so contention groups and Contest/Lottery resolvers see every contender |
| goal deadlines | scheduled as system events at init — a deadline fires at its exact timestamp under any policy; work completing exactly at the deadline counts |
| runner faults | a runner exception or decision_timeout_s (default 60 s wall) timeout becomes that agent's noop, recorded as runner_error: on the wake — one bad agent never stalls or kills the batch |
Messages: an accepted action whose schema has delivers=Delivery(...)
schedules per-recipient arrivals with class-based latency; inboxes are
consumed at observation time. Processes (Process(name, every_ms, fn)) run
world dynamics on the sim clock — never per-round — so the same scenario
behaves identically under every policy (pinned by
test_process_ticks_are_policy_independent).
Policies implement start / window_ms / on_batch_closed / on_agent_idle / on_action_complete / on_message and only ever call schedule_wake.
arun(pace=...) couples sim time to the wall clock for live spectating.
PerceptionSpec(sight, minimap, minimap_radius, hide_state, max_sight, max_messages, max_inventory) — sight is Chebyshev radius on grids,
route-hops elsewhere; 0 = own place only. State prefixed _ is never
observable. Size budgets bound every observation: overflow is summarized as
counts (Observation.omitted), so a crowded cell can't blow out a prompt.
The minimap draws terrain across its radius but live occupants only within
sight. Event visibility is witness-dependent: you
see your own outcomes, things that touched entities in your sight, and
messages addressed to you — sabotage at an empty bridge is genuinely unseen.
render_text is versioned (OBS_VERSION) and golden-tested
(test_golden_observation_text). Changing its wording is an experiment
change — bump the version and expect agent behavior to shift.
warren-trace/1, line-delimited JSON, identical on disk and over the live
WebSocket:
{"kind":"header", format, scenario, desc, seed, agents, npcs, sight, theme, meta}
{"kind":"genesis", world: {entities, relations}}
{"kind":"event", seq, t, actor, action, status, reason, args, delta}
{"kind":"wake", t, agent, reason, decision, outcome, text}
{"kind":"goals", t, goals: [...]} # emitted only on change
{"kind":"end", t, status, score, telemetry}
Invariant (pinned by test_trace_replay_reproduces_final_world): folding
every event delta over genesis reproduces the live final world exactly. The
viewer, scoring, and replay all rely on this — if you add a delta op, update
serialize_delta/deserialize_delta and the viewer's applyEvent together.
Event statuses: accepted, rejected, progress (plan step; reason: "started" marks submission), aborted (guard failed mid-plan), system
(processes, __deliver__).
warren/viz/assets/viewer.js (vanilla JS + vendored PixiJS, no build step).
State at time t = genesis folded forward, with snapshots every 400 events
for backscrub; layout comes from each place's Pos component (spaces in
authoring order); sprites/colors from the scenario theme
(id: > template: > kind: precedence, terrain: for tiles). Follow-mode
fog uses header.sight. Debug tip: everything is a global — open the console
and call setTime(...), setFollow('id'), inspect world.ents.
export_html inlines pixi + viewer + trace into one file; live_html points
the same viewer at a WebSocket. warren.serve broadcasts history-then-tail
per spectator.
Rendering resolves each key through three data layers, most specific first:
scenario skin pack (window.WARREN_SKIN_PACK, from s.skin(...) via
warren/viz/skin.py — animated 4-direction sprites, terrain variants,
multi-tile statics) → global pixel skin (window.WARREN_SKIN, built by
examples/build_skin.py) → emoji theme (trace header). Packs are pure
data (skin.toml + sheets, no code), so adding one can never change replay
semantics; scenarios/skins/overcooked_lite/ is the reference pack.
uv run pytest # whole suite, ~3 min (remote long-poll tests dominate)- Kernel/runtime: unit tests per contract (
test_core_*,test_runtime). - Scenarios are acceptance tests:
solution_runners()must complete the mission; distinctive mechanics get targeted tests. - Hardening (
test_hardening.py): chaosRandomRunners on every scenario with world invariants checked afterward (occupancy counts, single-holder, no dangling battles); trace-replay equivalence; same-process determinism; despawn-mid-plan; abort-frees-lane. Add a chaos invariant when you add a pack — the monkeys found the only real state-corruption bug so far. - Golden observation text; error messages are API (tests assert reasons).
Current scale: hundreds of agents, thousands of places, in-process. Known
headroom, in the order it will matter: (1) ReachablePlace.domain pathfinds
per POI per observation — cache per (place, epoch) if profiles say so;
(2) World.clone() per multi-effect action — replace with a COW overlay;
(3) trace wake records carry full observation text — gate behind a flag for
100k-agent runs; (4) viewer PIXI.Text labels — switch to BitmapText past a
few thousand entities.
Simulation.run()wrapsasyncio.run— inside an event loop useawait sim.arun(...).- Runner exceptions are not caught — a crashing runner crashes the run (deliberate in dev; wrap LLM runners in try/except like the example).
Spawn.makemust be pure (called for loci and again for the delta).- Components are frozen config: to change something at runtime it belongs in
state. - Movement/boarding require
OnFoot; scenario-local placement effects should respect occupancy counters or use the vehicles pack.