Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,9 @@ provider usage. Every agent is visible.

1. **Agent Zero plans through OpenRouter** on the first tick, periodic review,
or a material change such as directive completion, expiry, or player pressure.
Other ticks reuse the current directives without a planner call.
The server compiles bounded semantic strategic options per worker, and Agent
Zero selects an opaque option ID for each; the request carries no raw H3 cell
or agent IDs. Other ticks reuse the current directives without a planner call.
2. **Workers resolve directives through TypeSafe Jev** using compact observations
and opaque, engine-legal action candidates. Workers share a frozen pre-action
world; their calls currently run sequentially under one tick deadline.
Expand Down Expand Up @@ -87,7 +89,7 @@ are in [the screenshot guide](docs/assets/README.md).
- **Inspection:** switch between Live and Agents while the same execution
controller stays mounted. Inspect Zero strategy, worker directives, reflex
choices, validation outcomes, territory, and safe activity records.
- **Research exports:** generate compact or pretty schema-v12 JSON, download
- **Research exports:** generate compact or pretty schema-v13 JSON, download
it, or manually save the exact generated artifact to local SQLite. Exports
include bounded safe tick and provider-attempt records, including attempts
that did not produce a committed tick.
Expand Down Expand Up @@ -162,7 +164,7 @@ probe. Neither that probe nor `compare:live` runs in default tests or CI.
- [ADR 0033](docs/adr/0033-retire-legacy-multi-agent-architecture.md) — retirement
of the previous architecture

Current code reads only schema-v12 exports. Pre-swarm scenarios, snapshots, and
Current code reads only schema-v13 exports. Pre-swarm scenarios, snapshots, and
exports require an older Git revision. Historical ADRs and experiment reports
remain as decision history.

Expand Down
12 changes: 9 additions & 3 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,9 @@ architecture and removed all legacy infrastructure:
`longestRepeatedDirectionStreak`, `recentCellRevisits`), which nothing had
ever assigned since the migration, are computed per agent from accepted
moves.
- **PR #73** (`feat(swarm): compile semantic strategic options for Agent Zero`):
introduced `swarm-planner-v2` semantic strategic options and advanced export
schema to version 13. See ADR 0034.

The retirement case is structural rather than measured: the legacy path made one
full generative provider call per active agent per tick, so provider attempts,
Expand All @@ -57,7 +60,10 @@ implementation history.

Agent Zero is the sole generative planner. It makes an OpenRouter planning call
only when strategic replanning is required, regardless of roster size, under
the versioned contract `swarm-planner-v1`.
the versioned contract `swarm-planner-v2`. The server compiles bounded semantic
strategic options for each worker and Agent Zero selects an opaque option ID per
worker; the server resolves each option to a mission and target cell, so the
model request carries no raw H3 cell or agent IDs (see ADR 0034).
Each plan carries a strategy summary and one directive per active worker;
directives persist across ticks and are reused when no replan is triggered.
Agent Zero participates in the same frozen-world, simultaneous-tick transaction
Expand Down Expand Up @@ -165,14 +171,14 @@ paths or SQL, recovery, scheduling, MCP, and archive authority remain deferred.

Persistent short- and long-term objectives, compact memories, plan revision, summaries, and longer simulation runs.

_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 12. See ADR 0033._
_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 13. See ADRs 0033 and 0034._

## PR 6 — Persistent autonomous world

Scheduled turns, snapshots, replay, retries, idempotency, durable budget/attempt
ledgers, failure recovery, and operation without the World Lab browser being
open. The current process-local attempt and credit-admission ceilings are an
operator safety boundary, and their schema-v12 safe ledger can be exported to
operator safety boundary, and their schema-v13 safe ledger can be exported to
the analysis archive even when no turn committed. This is not active runtime
persistence, restart recovery, or provider-account balance enforcement.

Expand Down
32 changes: 16 additions & 16 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,13 +56,13 @@ World Lab distinguishes provider-reported cost from admission exposure. Each
attempt with unknown monetary cost, including TypeSafe Jev, retains its
configured per-attempt credit reserve in admission exposure. That reserve is
a conservative execution limit, not a measured charge or Jev cost estimate.
The OpenRouter planner asks Zero for bounded worker IDs, strategic target choice
IDs, mission, priority, risk, and its own legal action choice. For this compact
wire format, server code materializes agent IDs, H3 targets, directive IDs, and
five-tick lifetimes from the frozen observation. Full valid plans remain
accepted for compatibility, subject to validation that caps their lifetime at
ten ticks including the issue tick. Invalid output is classified into safe validation reasons
without retaining raw provider text. The selected Zero reasoning profile is
The OpenRouter planner asks Zero for one opaque `optionId` per offered worker
plus priority, risk, and its own legal action choice. For this compact wire
format, server code resolves each `optionId` to its mission and H3 target and
materializes agent IDs, directive IDs, and five-tick lifetimes from the frozen
observation; the model request contains no raw H3 cell IDs or agent IDs.
Invalid output is classified into safe validation reasons without retaining raw
provider text. The selected Zero reasoning profile is
sent to OpenRouter with private reasoning excluded, and output is bounded.

When Zero has no unexpired directive for a worker, a failed planning attempt
Expand Down Expand Up @@ -146,7 +146,7 @@ The Live workspace is a grid of independently scrolling agent rail, map, context

The Game API also owns one process-local experiment record. Each completed safe swarm tick is captured once, independently from the browser snapshot, and server-side export filters apply without affecting provider requests.

Schema-v12 exports may cross a separate offline archive boundary into `packages/experiment-archive`. Node's built-in SQLite stores normalized immutable research records through versioned migrations, foreign keys, prepared statements, and transactional idempotent imports. This downstream observability archive is never consulted by tick execution and cannot recover, resume, or mutate the active world. Its bounded query service is application-independent so a future read-only MCP adapter can reuse it without exposing arbitrary SQL.
Schema-v13 exports may cross a separate offline archive boundary into `packages/experiment-archive`. Node's built-in SQLite stores normalized immutable research records through versioned migrations, foreign keys, prepared statements, and transactional idempotent imports. This downstream observability archive is never consulted by tick execution and cannot recover, resume, or mutate the active world. Its bounded query service is application-independent so a future read-only MCP adapter can reuse it without exposing arbitrary SQL.

World Setup uses `world-scenario-v1`. Pure preview computes the actual H3 disk, exact count, summed cell area, deterministic roster/spawns, feasibility, and warnings. Apply recomputes and atomically replaces world and experiment state. Reset reconstructs the current scenario; the Toledo default preserves legacy starts. Explicit location search crosses a replaceable server-owned adapter with no autocomplete, a one-request-per-second Nominatim limit, bounded cache/timeout, normalized results, and OpenStreetMap attribution. Manual coordinates bypass that network boundary.

Expand Down Expand Up @@ -185,7 +185,7 @@ Equivalent legal moves are ordered reproducibly from world seed, stable agent ID
- `POST /api/simulation/experiment/export/archive` — import the exact generated safe document into the configured local SQLite archive
- `GET /api/simulation/models` — return the cached, sanitized compatible model catalog
- `POST /api/simulation/models/refresh` — explicitly refresh that catalog
- `POST /api/simulation/models/verify` — make one explicit, non-mutating compatibility probe against `swarm-planner-v1`
- `POST /api/simulation/models/verify` — make one explicit, non-mutating compatibility probe against `swarm-planner-v2`
- `POST /api/simulation/experiment/models` — replace the Agent Zero model assignment

The `GET /api/development-world` and `GET /health` endpoints remain for low-level diagnostics.
Expand Down Expand Up @@ -282,17 +282,17 @@ metrics cannot drift apart. Movement-pattern metrics walk each agent's accepted
moves separately, classifying each step with `geographicDirectionBetweenCells`;
aggregates sum direction counts and revisits and report the longest
single-agent streak. All exports
use schema version 12, which carries `swarmArchitectureVersion: "zero-swarm-v1"`
and independent provider-attempt accounting unconditionally. Pre-swarm exports
(schema versions 9, 10, and 11) are rejected outright; there is no migration
path. The provider-attempt ledger is canonical for attempt counts, latency,
token, and cost totals.
use schema version 13, which carries `swarmArchitectureVersion: "zero-swarm-v1"`
and independent provider-attempt accounting unconditionally. Exports at schema
version 12 and earlier are rejected outright; there is no migration path. The
provider-attempt ledger is canonical for attempt counts, latency, token, and
cost totals.

The agent runtime follows [OpenRouter's usage-accounting contract](https://openrouter.ai/docs/cookbook/administration/usage-accounting) and normalizes optional non-streaming usage fields: prompt, completion, total, reasoning, cached-read, cache-write tokens, and actual `usage.cost` as `costCredits`. It never derives price from a table. Safe usage already returned with a billable response is retained on later decision JSON/schema failure; network and HTTP failures without usage remain unknown. Scripted providers explicitly report zero tokens and zero cost.

## Packages

`packages/shared` owns centralized scenario limits and all public schemas, including model capabilities, swarm directives, metrics, and schema-v12 swarm tick exports. Other-agent observations remain deterministically capped at seven for larger rosters. Types are inferred from Zod.
`packages/shared` owns centralized scenario limits and all public schemas, including model capabilities, swarm directives, metrics, and schema-v13 swarm tick exports. Other-agent observations remain deterministically capped at seven for larger rosters. Types are inferred from Zod.

`packages/world-engine` remains deterministic and has no model, HTTP, UI, storage, or credential dependency. It validates world actions independently. Direct proximity is derived from a separately supplied pre-action state.

Expand Down Expand Up @@ -327,5 +327,5 @@ Structural provider failures retain the broad compatibility code plus bounded de

## Provider-attempt accounting

Provider work has an independent bounded lifecycle ledger. Schema-v12 exports
Provider work has an independent bounded lifecycle ledger. Schema-v13 exports
and archive-v4 preserve safe attempt records even when no world tick commits.
14 changes: 8 additions & 6 deletions docs/EXPERIMENT_ARCHIVE.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
# Local experiment archive

The archive accepts only schema-v12 exports and rejects any other schema version outright.
Pre-swarm exports (schema versions 9, 10, and 11) are rejected with no migration path; reading them
requires checking out the Git revision before PR 1 of the zero-swarm migration.
The archive accepts only schema-v13 exports and rejects any other schema version outright.
Exports at schema version 12 and earlier are rejected with no migration path. Reading a
schema-v12 export requires a Git revision before PR #73 (`swarm-planner-v2`); reading a
pre-swarm export (schema versions 9, 10, and 11) requires a revision before PR 1 of the
zero-swarm migration.
Migration 2 adds nullable tick number, deterministic tick position, virtual
time, and interval columns. Bounded queries order tick-attributed records by
tick and tick position where exposed; the CLI still provides no arbitrary SQL
Expand All @@ -21,7 +23,7 @@ Migration 6 removes all legacy per-agent-LLM social-system tables: `turns`,
`turn_number` column is renamed `tick_number`. Personality and behavior columns
are dropped from `agents` and `experiments`.

The experiment archive is a durable, local research surface for completed or partially retained exports. It does not participate in an active simulation: the Game API's in-memory engine remains authoritative, and an archive write cannot change an accepted game outcome. It imports schema-v12 JSON exports only; it is not crash recovery, restartable simulation state, or a scheduler.
The experiment archive is a durable, local research surface for completed or partially retained exports. It does not participate in an active simulation: the Game API's in-memory engine remains authoritative, and an archive write cannot change an accepted game outcome. It imports schema-v13 JSON exports only; it is not crash recovery, restartable simulation state, or a scheduler.

## Storage and configuration

Expand Down Expand Up @@ -102,7 +104,7 @@ MCP and embeddings are deferred because bounded local retrieval solves the immed
Archive schema v4 stores `providerAttempts` independently. Use
`pnpm experiment:db provider-attempts <experiment-id>` to inspect committed and
uncommitted provider work. Monetary values round-trip as canonical TEXT. This
ledger is canonical for every current (schema-v12) export, which always
ledger is canonical for every current (schema-v13) export, which always
carries independent attempt accounting. The SQLite archive is for analysis
and is not active runtime recovery.

Expand All @@ -114,7 +116,7 @@ swarm ticks.

## Zero-swarm comparisons

The archive preserves safe schema-v12 swarm tick records and independent
The archive preserves safe schema-v13 swarm tick records and independent
provider attempts, but its `compare` command is not the same-scenario,
per-tick swarm harness. Use `pnpm compare:offline` for the reproducible
zero-swarm-vs-deterministic-workers fixture report. The runner does not
Expand Down
9 changes: 6 additions & 3 deletions docs/GAMEPLAY_FOUNDATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,11 +21,14 @@ without a planner call. Agent Zero issues a strategy summary and one structured
directive per worker. Workers resolve their directive via TypeSafe Jev reflex
cognition over enumerated legal-action candidates; deterministic selection is
the fallback when Jev is unavailable or its output fails validation. Agent
Zero is also a roster agent and selects its own legal physical action. The
server assigns directive IDs, target cells, and bounded lifetimes.
Zero is also a roster agent and selects its own legal physical action. Under
`swarm-planner-v2` the server compiles bounded semantic strategic options per
worker and Agent Zero selects an opaque option ID for each; the server resolves
it to a mission and target cell and assigns directive IDs and bounded lifetimes.
The model request carries no raw H3 cell or agent IDs.
Current setup rejects a missing or null Agent Zero designation. Pre-swarm
exports (schema versions 9–11) are not readable by current code; schema
version 12 with `swarmArchitectureVersion: 'zero-swarm-v1'` is required.
version 13 with `swarmArchitectureVersion: 'zero-swarm-v1'` is required.

> **Status: accepted product and roadmap direction, not an implementation claim.**
> This document records foundational decisions for future World Lab and Player
Expand Down
4 changes: 2 additions & 2 deletions docs/SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,7 @@ observation data. Raw planner text is not forwarded. Workers do not receive
other workers' choices. Cancellation discards every result from the
uncommitted tick.

When strategic replanning is required, the OpenRouter planner receives one bounded strategic observation and is instructed to return exactly one plain JSON object naming opaque worker and target choices plus a Zero-action selection. TypeSafe Jev receives a compact semantic observation with opaque legal candidate IDs per worker and returns a probability distribution over candidates; a second question in the same request returns an optional bounded replan probability. The runtime performs bounded extraction and conservative repair for wrappers such as code fences, surrounding prose, and trailing commas, then rejects missing text, unusable JSON, unknown fields, or output truncation before the deterministic world engine validates all resolved components independently.
When strategic replanning is required, the OpenRouter planner receives one bounded strategic observation (a coarse `worldSummary` plus per-worker semantic `workerOptions`, with no raw H3 cell IDs or agent IDs) and is instructed to return exactly one plain JSON object naming opaque worker and option choices plus a Zero-action selection. TypeSafe Jev receives a compact semantic observation with opaque legal candidate IDs per worker and returns a probability distribution over candidates; a second question in the same request returns an optional bounded replan probability. The runtime performs bounded extraction and conservative repair for wrappers such as code fences, surrounding prose, and trailing commas, then rejects missing text, unusable JSON, unknown fields, or output truncation before the deterministic world engine validates all resolved components independently.

The request uses the selected model, messages, `max_tokens`, `stream: false`, and at most one normalized reasoning object selected from sanitized model metadata. Provider default omits the object. Off is offered only for non-mandatory reasoning and sends `{ enabled: false, exclude: true }`; an advertised effort sends `{ enabled: true, effort, exclude: true }`. It deliberately sends no tools, `tool_choice`, `response_format`, `provider.require_parameters`, standalone `reasoning_effort`, or model-specific parameter. Model IDs are never inspected or special-cased. Transport/provider failures, unavailable-model/profile failures, text/JSON contract failures, and later simulation-rule rejection remain distinct safe outcomes. The adapter never silently substitutes a model or scripted behavior.

Expand Down Expand Up @@ -155,7 +155,7 @@ The Game API captures only schema-validated safe observations, requested world a

Export requests, agent IDs, levels, ranges, and Custom dependencies are
runtime-validated. Filtering and metrics remain server-owned. The export schema
is exclusively version 12; exports carrying schema version 9, 10, or 11 are
is exclusively version 13; exports carrying schema version 12 or earlier are
rejected outright with no migration path. Reset clears swarm tick history and
metrics while unlocking preserved roster assignments for the new experiment.

Expand Down
4 changes: 2 additions & 2 deletions docs/TESTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,9 +79,9 @@ malformed-entry skipping, cache TTL, stale fallback, and safe failure states.

### Experiment archive (`packages/experiment-archive`)

`archive.test.ts` covers schema-v12-only enforcement: swarm-native provenance
`archive.test.ts` covers schema-v13-only enforcement: swarm-native provenance
archival, idempotent import, query service, credential-like-data rejection before
persistence, unknown-architecture-version rejection, and non-v12 schema
persistence, unknown-architecture-version rejection, and non-v13 schema
rejection.

### Game API (`apps/game-api`)
Expand Down
2 changes: 1 addition & 1 deletion docs/ZERO_SWARM_COMPARISON.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ reproducibility and exposes comparable telemetry; it does not demonstrate
real-model quality, production latency, or provider cost. Real-provider studies
remain explicitly opted in and should archive their safe exports separately.

The experiment archive importer now accepts only schema-v12 swarm-native
The experiment archive importer now accepts only schema-v13 swarm-native
exports, which always carry swarm tick records and independent provider
attempt accounting; it does not replace this same-scenario, per-tick swarm
harness.
Expand Down
Loading