Skip to content

Latest commit

 

History

History
265 lines (212 loc) · 13.7 KB

File metadata and controls

265 lines (212 loc) · 13.7 KB

WARREN Developer Guide

The engine's internals, its contracts, and where to extend it. Companions: DESIGN.md (why it's shaped this way), authoring.md (writing scenarios), and the test suite — every contract described here is pinned by a test in tests/.

Layer map

warren.core        world · components · params DSL · effects · adjudication
warren.runtime     sim clock · wake policies · durative plans · messages · processes
warren.space       grid/room compilers · pathfinding · move / move_to
warren.perception  visibility · observations · versioned text renderer · themes
warren.mechanics   packs (inventory, resources, comms, damage, vehicles, production, encounters)
warren.author      scenario micro-DSL + build-time validation
warren.agents      runners (scripted / random / policy)
warren.viz         trace recorder · PixiJS viewer
warren.protocol    warren-agent/0 wire models (no engine imports)
warren.remote      agent mailboxes · RemoteRunner · briefings/manifest builders
warren.serve       FastAPI host (spectators + /v1 agent control plane)
warren.client      httpx-only protocol client + framework adapters (pydantic-ai)
warren.analysis    scenario-agnostic trace summaries

Dependencies point strictly downward. The kernel knows nothing about time policy, perception, or geometry; mechanics packs are ordinary scenario-level code that happens to ship in the library.

Execution lifecycle

WARREN event loop

The heap is the only clock. When its head is a system event (durative completion, plan step, message arrival, process tick, goal deadline) the runtime drains everything due at that instant straight into adjudication — no agent is consulted. When the head is a wake, it closes a batch that never extends past the next system event, builds every observation against the same world, gathers decisions concurrently, and applies them in agent order. Both paths go through the same deterministic engine.step, whose events feed the wake policy (who wakes next), goal evaluation, and the trace. A fuller, per-decision-branch version of this figure lives in media/lifecycle.mmd (SVG).

The kernel (warren.core)

World model

Entity(id, kind, template, name, desc, tags, components, state) — the only mutable field is state. components holds frozen dataclasses registered via @component("name"); they serialize by registry name, so import the module that defines a component before deserializing worlds that use it.

Relations (src, kind, dst, plus props/state) carry all structure: at (placement), holds (inventory), route (topology, with state.open/state.blocked_by), and any scenario-defined kinds. The world maintains indexes — by kind, by component, occupants-by-container, relation adjacency both ways — so perception and affordances never scan all entities.

place_of() follows the containment chain (held item → holder → vehicle → cell); container_of() is one hop. Mutation happens only through World.apply(delta).

Params: the single binding layer

The design's core move. A param spec answers three questions from one declaration:

params={"item": EntityParam(comp="carryable", scope="colocated")}
  1. validate a submitted arg (wrong kind/scope/filter → teaching rejection like item:not_here:gem),
  2. enumerate the candidate domain for observations (from indexes, capped),
  3. schema for LLM tool calling.

ActionSchema.tool_view() builds the agent-facing tool: per-param choices for finite domains, plus fully-validated concrete suggestions when the combination space is small (≤200 combos, capped at 10 suggestions). A required finite param with an empty domain hides the tool. Free-form params (IntParam, StrParam) surface as schema only.

When adding a param type, implement domain / validate / schema and keep the invariant: everything domain returns must pass validate.

Effects and write-loci

Declarative effects (Set, Inc, Take, PlaceAt, Transfer, Spawn, …) compute a Delta and report the loci they write. Custom Fx(fn, at=...) must declare loci or it conservatively locks its actor. Loci shapes: ("state", id, field), ("place", id), ("own", id), ("rel", src, kind, dst, field), ("id", id).

Multi-effect actions resolve against a scratch clone so a failing effect can't half-apply (see ActionSchema.resolve — swap in a COW overlay if world clones ever show up in profiles).

Adjudication

Engine.step(actions, time):

  1. reject duplicates / unknowns; validate params
  2. compute loci; union-find proposals whose loci overlap into groups
  3. per group: schema-owned GroupResolvers adjudicate jointly (e.g. Contest); everything else re-validates sequentially in actor order against the updated world — which is why a depleted well or a taken item rejects later contenders with an honest reason, and a capacity-2 train admits exactly 2 of 3 simultaneous boarders
  4. apply deltas, append Events to the log

Multi-effect actions resolve to a canonical net delta (relations as a presence diff, mixed set/inc collapsed to final values): the authored effect sequence is honored and replay matches live execution. World.apply preflights every reference — a bad delta raises DeltaError before any mutation, which the engine converts into a rejected event.

Determinism contract: world transitions are a pure function of (genesis, ordered action stream, seed). Every RNG derives from derive_rng(seed, *key). Scheduling policies change which actions get proposed — histories differ across policies by design; tests pin the former, never the latter. Corollaries: never call random/hash() in mechanics or policies (use ctx.rng / zlib.crc32), and never mint ids from global counters (see the battle-id regression test).

The runtime (warren.runtime)

One heap queue in simulated ms. System entries (plan steps, deliveries, process ticks) sort before wakes at the same instant, and a wake batch never extends past the next system event (conservative causality: everyone in a batch decides on the same world).

The loop per batch: collect wakes in the window → build all observations against the same state → gather decisions (concurrently — async runners overlap within a batch) → apply in deterministic agent order. Caveat: in busy worlds batches are near-singletons (measured mean 1.54 agents on the 100-agent city run), so slow runners still serialize across batches. The planned fix — treating LLM thinking time as a simulated duration the scheduler can look ahead through — is on the roadmap.

Decisions: Act(name, args) | Sleep(ms) | None. What happens next:

Case Behavior
instantaneous action batched into one engine.step at batch time
duration= action validated now for fast feedback; adjudicated again at completion via engine.step (so completion-time contention is correct)
plan= action validated now; steps execute on the clock, each re-checking its guard; failure emits aborted and frees the lane
busy lane rejection + re-wake when the lane frees (lane_free), not a retry spin
wake while busy observation carries busy (lane, action, args, until_ms); tools excludes busy-lane actions, so the agent is never offered an action it can only get rejected for
rejection policy re-asks after rejection_ms (default 400ms), not 1ms
Sleep(ms) swallows all wakes except messages until the deadline; each completion wake is stamped with its deadline, so re-sleeping after a message discards the superseded wake (no accumulation) while the current one still fires even when batching pulls it in early
simultaneous completions default-durative actions finishing at the same instant adjudicate as one Engine.step cohort, so contention groups and Contest/Lottery resolvers see every contender
goal deadlines scheduled as system events at init — a deadline fires at its exact timestamp under any policy; work completing exactly at the deadline counts
runner faults a runner exception or decision_timeout_s (default 60 s wall) timeout becomes that agent's noop, recorded as runner_error: on the wake — one bad agent never stalls or kills the batch

Messages: an accepted action whose schema has delivers=Delivery(...) schedules per-recipient arrivals with class-based latency; inboxes are consumed at observation time. Processes (Process(name, every_ms, fn)) run world dynamics on the sim clock — never per-round — so the same scenario behaves identically under every policy (pinned by test_process_ticks_are_policy_independent).

Policies implement start / window_ms / on_batch_closed / on_agent_idle / on_action_complete / on_message and only ever call schedule_wake.

arun(pace=...) couples sim time to the wall clock for live spectating.

Perception (warren.perception)

PerceptionSpec(sight, minimap, minimap_radius, hide_state, max_sight, max_messages, max_inventory) — sight is Chebyshev radius on grids, route-hops elsewhere; 0 = own place only. State prefixed _ is never observable. Size budgets bound every observation: overflow is summarized as counts (Observation.omitted), so a crowded cell can't blow out a prompt. The minimap draws terrain across its radius but live occupants only within sight. Event visibility is witness-dependent: you see your own outcomes, things that touched entities in your sight, and messages addressed to you — sabotage at an empty bridge is genuinely unseen.

render_text is versioned (OBS_VERSION) and golden-tested (test_golden_observation_text). Changing its wording is an experiment change — bump the version and expect agent behavior to shift.

Trace format (warren.viz)

warren-trace/1, line-delimited JSON, identical on disk and over the live WebSocket:

{"kind":"header", format, scenario, desc, seed, agents, npcs, sight, theme, meta}
{"kind":"genesis", world: {entities, relations}}
{"kind":"event",  seq, t, actor, action, status, reason, args, delta}
{"kind":"wake",   t, agent, reason, decision, outcome, text}
{"kind":"goals",  t, goals: [...]}          # emitted only on change
{"kind":"end",    t, status, score, telemetry}

Invariant (pinned by test_trace_replay_reproduces_final_world): folding every event delta over genesis reproduces the live final world exactly. The viewer, scoring, and replay all rely on this — if you add a delta op, update serialize_delta/deserialize_delta and the viewer's applyEvent together.

Event statuses: accepted, rejected, progress (plan step; reason: "started" marks submission), aborted (guard failed mid-plan), system (processes, __deliver__).

Viewer

warren/viz/assets/viewer.js (vanilla JS + vendored PixiJS, no build step). State at time t = genesis folded forward, with snapshots every 400 events for backscrub; layout comes from each place's Pos component (spaces in authoring order); sprites/colors from the scenario theme (id: > template: > kind: precedence, terrain: for tiles). Follow-mode fog uses header.sight. Debug tip: everything is a global — open the console and call setTime(...), setFollow('id'), inspect world.ents.

export_html inlines pixi + viewer + trace into one file; live_html points the same viewer at a WebSocket. warren.serve broadcasts history-then-tail per spectator.

Rendering resolves each key through three data layers, most specific first: scenario skin pack (window.WARREN_SKIN_PACK, from s.skin(...) via warren/viz/skin.py — animated 4-direction sprites, terrain variants, multi-tile statics) → global pixel skin (window.WARREN_SKIN, built by examples/build_skin.py) → emoji theme (trace header). Packs are pure data (skin.toml + sheets, no code), so adding one can never change replay semantics; scenarios/skins/overcooked_lite/ is the reference pack.

Testing conventions

uv run pytest              # whole suite, ~3 min (remote long-poll tests dominate)
  • Kernel/runtime: unit tests per contract (test_core_*, test_runtime).
  • Scenarios are acceptance tests: solution_runners() must complete the mission; distinctive mechanics get targeted tests.
  • Hardening (test_hardening.py): chaos RandomRunners on every scenario with world invariants checked afterward (occupancy counts, single-holder, no dangling battles); trace-replay equivalence; same-process determinism; despawn-mid-plan; abort-frees-lane. Add a chaos invariant when you add a pack — the monkeys found the only real state-corruption bug so far.
  • Golden observation text; error messages are API (tests assert reasons).

Performance notes

Current scale: hundreds of agents, thousands of places, in-process. Known headroom, in the order it will matter: (1) ReachablePlace.domain pathfinds per POI per observation — cache per (place, epoch) if profiles say so; (2) World.clone() per multi-effect action — replace with a COW overlay; (3) trace wake records carry full observation text — gate behind a flag for 100k-agent runs; (4) viewer PIXI.Text labels — switch to BitmapText past a few thousand entities.

Gotchas

  • Simulation.run() wraps asyncio.run — inside an event loop use await sim.arun(...).
  • Runner exceptions are not caught — a crashing runner crashes the run (deliberate in dev; wrap LLM runners in try/except like the example).
  • Spawn.make must be pure (called for loci and again for the delta).
  • Components are frozen config: to change something at runtime it belongs in state.
  • Movement/boarding require OnFoot; scenario-local placement effects should respect occupancy counters or use the vehicles pack.