diff --git a/README.md b/README.md index 2a63e30..6fa2a1b 100644 --- a/README.md +++ b/README.md @@ -18,7 +18,9 @@ provider usage. Every agent is visible. 1. **Agent Zero plans through OpenRouter** on the first tick, periodic review, or a material change such as directive completion, expiry, or player pressure. - Other ticks reuse the current directives without a planner call. + The server compiles bounded semantic strategic options per worker, and Agent + Zero selects an opaque option ID for each; the request carries no raw H3 cell + or agent IDs. Other ticks reuse the current directives without a planner call. 2. **Workers resolve directives through TypeSafe Jev** using compact observations and opaque, engine-legal action candidates. Workers share a frozen pre-action world; their calls currently run sequentially under one tick deadline. @@ -87,7 +89,7 @@ are in [the screenshot guide](docs/assets/README.md). - **Inspection:** switch between Live and Agents while the same execution controller stays mounted. Inspect Zero strategy, worker directives, reflex choices, validation outcomes, territory, and safe activity records. -- **Research exports:** generate compact or pretty schema-v12 JSON, download +- **Research exports:** generate compact or pretty schema-v13 JSON, download it, or manually save the exact generated artifact to local SQLite. Exports include bounded safe tick and provider-attempt records, including attempts that did not produce a committed tick. @@ -162,7 +164,7 @@ probe. Neither that probe nor `compare:live` runs in default tests or CI. - [ADR 0033](docs/adr/0033-retire-legacy-multi-agent-architecture.md) — retirement of the previous architecture -Current code reads only schema-v12 exports. Pre-swarm scenarios, snapshots, and +Current code reads only schema-v13 exports. Pre-swarm scenarios, snapshots, and exports require an older Git revision. Historical ADRs and experiment reports remain as decision history. diff --git a/ROADMAP.md b/ROADMAP.md index 7fa6503..9507e75 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -38,6 +38,9 @@ architecture and removed all legacy infrastructure: `longestRepeatedDirectionStreak`, `recentCellRevisits`), which nothing had ever assigned since the migration, are computed per agent from accepted moves. +- **PR #73** (`feat(swarm): compile semantic strategic options for Agent Zero`): + introduced `swarm-planner-v2` semantic strategic options and advanced export + schema to version 13. See ADR 0034. The retirement case is structural rather than measured: the legacy path made one full generative provider call per active agent per tick, so provider attempts, @@ -57,7 +60,10 @@ implementation history. Agent Zero is the sole generative planner. It makes an OpenRouter planning call only when strategic replanning is required, regardless of roster size, under -the versioned contract `swarm-planner-v1`. +the versioned contract `swarm-planner-v2`. The server compiles bounded semantic +strategic options for each worker and Agent Zero selects an opaque option ID per +worker; the server resolves each option to a mission and target cell, so the +model request carries no raw H3 cell or agent IDs (see ADR 0034). Each plan carries a strategy summary and one directive per active worker; directives persist across ticks and are reused when no replan is triggered. Agent Zero participates in the same frozen-world, simultaneous-tick transaction @@ -165,14 +171,14 @@ paths or SQL, recovery, scheduling, MCP, and archive authority remain deferred. Persistent short- and long-term objectives, compact memories, plan revision, summaries, and longer simulation runs. -_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 12. See ADR 0033._ +_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 13. See ADRs 0033 and 0034._ ## PR 6 — Persistent autonomous world Scheduled turns, snapshots, replay, retries, idempotency, durable budget/attempt ledgers, failure recovery, and operation without the World Lab browser being open. The current process-local attempt and credit-admission ceilings are an -operator safety boundary, and their schema-v12 safe ledger can be exported to +operator safety boundary, and their schema-v13 safe ledger can be exported to the analysis archive even when no turn committed. This is not active runtime persistence, restart recovery, or provider-account balance enforcement. diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 6819fd0..7c5d0d8 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -56,13 +56,13 @@ World Lab distinguishes provider-reported cost from admission exposure. Each attempt with unknown monetary cost, including TypeSafe Jev, retains its configured per-attempt credit reserve in admission exposure. That reserve is a conservative execution limit, not a measured charge or Jev cost estimate. -The OpenRouter planner asks Zero for bounded worker IDs, strategic target choice -IDs, mission, priority, risk, and its own legal action choice. For this compact -wire format, server code materializes agent IDs, H3 targets, directive IDs, and -five-tick lifetimes from the frozen observation. Full valid plans remain -accepted for compatibility, subject to validation that caps their lifetime at -ten ticks including the issue tick. Invalid output is classified into safe validation reasons -without retaining raw provider text. The selected Zero reasoning profile is +The OpenRouter planner asks Zero for one opaque `optionId` per offered worker +plus priority, risk, and its own legal action choice. For this compact wire +format, server code resolves each `optionId` to its mission and H3 target and +materializes agent IDs, directive IDs, and five-tick lifetimes from the frozen +observation; the model request contains no raw H3 cell IDs or agent IDs. +Invalid output is classified into safe validation reasons without retaining raw +provider text. The selected Zero reasoning profile is sent to OpenRouter with private reasoning excluded, and output is bounded. When Zero has no unexpired directive for a worker, a failed planning attempt @@ -146,7 +146,7 @@ The Live workspace is a grid of independently scrolling agent rail, map, context The Game API also owns one process-local experiment record. Each completed safe swarm tick is captured once, independently from the browser snapshot, and server-side export filters apply without affecting provider requests. -Schema-v12 exports may cross a separate offline archive boundary into `packages/experiment-archive`. Node's built-in SQLite stores normalized immutable research records through versioned migrations, foreign keys, prepared statements, and transactional idempotent imports. This downstream observability archive is never consulted by tick execution and cannot recover, resume, or mutate the active world. Its bounded query service is application-independent so a future read-only MCP adapter can reuse it without exposing arbitrary SQL. +Schema-v13 exports may cross a separate offline archive boundary into `packages/experiment-archive`. Node's built-in SQLite stores normalized immutable research records through versioned migrations, foreign keys, prepared statements, and transactional idempotent imports. This downstream observability archive is never consulted by tick execution and cannot recover, resume, or mutate the active world. Its bounded query service is application-independent so a future read-only MCP adapter can reuse it without exposing arbitrary SQL. World Setup uses `world-scenario-v1`. Pure preview computes the actual H3 disk, exact count, summed cell area, deterministic roster/spawns, feasibility, and warnings. Apply recomputes and atomically replaces world and experiment state. Reset reconstructs the current scenario; the Toledo default preserves legacy starts. Explicit location search crosses a replaceable server-owned adapter with no autocomplete, a one-request-per-second Nominatim limit, bounded cache/timeout, normalized results, and OpenStreetMap attribution. Manual coordinates bypass that network boundary. @@ -185,7 +185,7 @@ Equivalent legal moves are ordered reproducibly from world seed, stable agent ID - `POST /api/simulation/experiment/export/archive` — import the exact generated safe document into the configured local SQLite archive - `GET /api/simulation/models` — return the cached, sanitized compatible model catalog - `POST /api/simulation/models/refresh` — explicitly refresh that catalog -- `POST /api/simulation/models/verify` — make one explicit, non-mutating compatibility probe against `swarm-planner-v1` +- `POST /api/simulation/models/verify` — make one explicit, non-mutating compatibility probe against `swarm-planner-v2` - `POST /api/simulation/experiment/models` — replace the Agent Zero model assignment The `GET /api/development-world` and `GET /health` endpoints remain for low-level diagnostics. @@ -282,17 +282,17 @@ metrics cannot drift apart. Movement-pattern metrics walk each agent's accepted moves separately, classifying each step with `geographicDirectionBetweenCells`; aggregates sum direction counts and revisits and report the longest single-agent streak. All exports -use schema version 12, which carries `swarmArchitectureVersion: "zero-swarm-v1"` -and independent provider-attempt accounting unconditionally. Pre-swarm exports -(schema versions 9, 10, and 11) are rejected outright; there is no migration -path. The provider-attempt ledger is canonical for attempt counts, latency, -token, and cost totals. +use schema version 13, which carries `swarmArchitectureVersion: "zero-swarm-v1"` +and independent provider-attempt accounting unconditionally. Exports at schema +version 12 and earlier are rejected outright; there is no migration path. The +provider-attempt ledger is canonical for attempt counts, latency, token, and +cost totals. The agent runtime follows [OpenRouter's usage-accounting contract](https://openrouter.ai/docs/cookbook/administration/usage-accounting) and normalizes optional non-streaming usage fields: prompt, completion, total, reasoning, cached-read, cache-write tokens, and actual `usage.cost` as `costCredits`. It never derives price from a table. Safe usage already returned with a billable response is retained on later decision JSON/schema failure; network and HTTP failures without usage remain unknown. Scripted providers explicitly report zero tokens and zero cost. ## Packages -`packages/shared` owns centralized scenario limits and all public schemas, including model capabilities, swarm directives, metrics, and schema-v12 swarm tick exports. Other-agent observations remain deterministically capped at seven for larger rosters. Types are inferred from Zod. +`packages/shared` owns centralized scenario limits and all public schemas, including model capabilities, swarm directives, metrics, and schema-v13 swarm tick exports. Other-agent observations remain deterministically capped at seven for larger rosters. Types are inferred from Zod. `packages/world-engine` remains deterministic and has no model, HTTP, UI, storage, or credential dependency. It validates world actions independently. Direct proximity is derived from a separately supplied pre-action state. @@ -327,5 +327,5 @@ Structural provider failures retain the broad compatibility code plus bounded de ## Provider-attempt accounting -Provider work has an independent bounded lifecycle ledger. Schema-v12 exports +Provider work has an independent bounded lifecycle ledger. Schema-v13 exports and archive-v4 preserve safe attempt records even when no world tick commits. diff --git a/docs/EXPERIMENT_ARCHIVE.md b/docs/EXPERIMENT_ARCHIVE.md index 6f312ac..063f567 100644 --- a/docs/EXPERIMENT_ARCHIVE.md +++ b/docs/EXPERIMENT_ARCHIVE.md @@ -1,8 +1,10 @@ # Local experiment archive -The archive accepts only schema-v12 exports and rejects any other schema version outright. -Pre-swarm exports (schema versions 9, 10, and 11) are rejected with no migration path; reading them -requires checking out the Git revision before PR 1 of the zero-swarm migration. +The archive accepts only schema-v13 exports and rejects any other schema version outright. +Exports at schema version 12 and earlier are rejected with no migration path. Reading a +schema-v12 export requires a Git revision before PR #73 (`swarm-planner-v2`); reading a +pre-swarm export (schema versions 9, 10, and 11) requires a revision before PR 1 of the +zero-swarm migration. Migration 2 adds nullable tick number, deterministic tick position, virtual time, and interval columns. Bounded queries order tick-attributed records by tick and tick position where exposed; the CLI still provides no arbitrary SQL @@ -21,7 +23,7 @@ Migration 6 removes all legacy per-agent-LLM social-system tables: `turns`, `turn_number` column is renamed `tick_number`. Personality and behavior columns are dropped from `agents` and `experiments`. -The experiment archive is a durable, local research surface for completed or partially retained exports. It does not participate in an active simulation: the Game API's in-memory engine remains authoritative, and an archive write cannot change an accepted game outcome. It imports schema-v12 JSON exports only; it is not crash recovery, restartable simulation state, or a scheduler. +The experiment archive is a durable, local research surface for completed or partially retained exports. It does not participate in an active simulation: the Game API's in-memory engine remains authoritative, and an archive write cannot change an accepted game outcome. It imports schema-v13 JSON exports only; it is not crash recovery, restartable simulation state, or a scheduler. ## Storage and configuration @@ -102,7 +104,7 @@ MCP and embeddings are deferred because bounded local retrieval solves the immed Archive schema v4 stores `providerAttempts` independently. Use `pnpm experiment:db provider-attempts ` to inspect committed and uncommitted provider work. Monetary values round-trip as canonical TEXT. This -ledger is canonical for every current (schema-v12) export, which always +ledger is canonical for every current (schema-v13) export, which always carries independent attempt accounting. The SQLite archive is for analysis and is not active runtime recovery. @@ -114,7 +116,7 @@ swarm ticks. ## Zero-swarm comparisons -The archive preserves safe schema-v12 swarm tick records and independent +The archive preserves safe schema-v13 swarm tick records and independent provider attempts, but its `compare` command is not the same-scenario, per-tick swarm harness. Use `pnpm compare:offline` for the reproducible zero-swarm-vs-deterministic-workers fixture report. The runner does not diff --git a/docs/GAMEPLAY_FOUNDATION.md b/docs/GAMEPLAY_FOUNDATION.md index 18c8211..511d0b0 100644 --- a/docs/GAMEPLAY_FOUNDATION.md +++ b/docs/GAMEPLAY_FOUNDATION.md @@ -21,11 +21,14 @@ without a planner call. Agent Zero issues a strategy summary and one structured directive per worker. Workers resolve their directive via TypeSafe Jev reflex cognition over enumerated legal-action candidates; deterministic selection is the fallback when Jev is unavailable or its output fails validation. Agent -Zero is also a roster agent and selects its own legal physical action. The -server assigns directive IDs, target cells, and bounded lifetimes. +Zero is also a roster agent and selects its own legal physical action. Under +`swarm-planner-v2` the server compiles bounded semantic strategic options per +worker and Agent Zero selects an opaque option ID for each; the server resolves +it to a mission and target cell and assigns directive IDs and bounded lifetimes. +The model request carries no raw H3 cell or agent IDs. Current setup rejects a missing or null Agent Zero designation. Pre-swarm exports (schema versions 9–11) are not readable by current code; schema -version 12 with `swarmArchitectureVersion: 'zero-swarm-v1'` is required. +version 13 with `swarmArchitectureVersion: 'zero-swarm-v1'` is required. > **Status: accepted product and roadmap direction, not an implementation claim.** > This document records foundational decisions for future World Lab and Player diff --git a/docs/SECURITY.md b/docs/SECURITY.md index 228bfa9..831b493 100644 --- a/docs/SECURITY.md +++ b/docs/SECURITY.md @@ -99,7 +99,7 @@ observation data. Raw planner text is not forwarded. Workers do not receive other workers' choices. Cancellation discards every result from the uncommitted tick. -When strategic replanning is required, the OpenRouter planner receives one bounded strategic observation and is instructed to return exactly one plain JSON object naming opaque worker and target choices plus a Zero-action selection. TypeSafe Jev receives a compact semantic observation with opaque legal candidate IDs per worker and returns a probability distribution over candidates; a second question in the same request returns an optional bounded replan probability. The runtime performs bounded extraction and conservative repair for wrappers such as code fences, surrounding prose, and trailing commas, then rejects missing text, unusable JSON, unknown fields, or output truncation before the deterministic world engine validates all resolved components independently. +When strategic replanning is required, the OpenRouter planner receives one bounded strategic observation (a coarse `worldSummary` plus per-worker semantic `workerOptions`, with no raw H3 cell IDs or agent IDs) and is instructed to return exactly one plain JSON object naming opaque worker and option choices plus a Zero-action selection. TypeSafe Jev receives a compact semantic observation with opaque legal candidate IDs per worker and returns a probability distribution over candidates; a second question in the same request returns an optional bounded replan probability. The runtime performs bounded extraction and conservative repair for wrappers such as code fences, surrounding prose, and trailing commas, then rejects missing text, unusable JSON, unknown fields, or output truncation before the deterministic world engine validates all resolved components independently. The request uses the selected model, messages, `max_tokens`, `stream: false`, and at most one normalized reasoning object selected from sanitized model metadata. Provider default omits the object. Off is offered only for non-mandatory reasoning and sends `{ enabled: false, exclude: true }`; an advertised effort sends `{ enabled: true, effort, exclude: true }`. It deliberately sends no tools, `tool_choice`, `response_format`, `provider.require_parameters`, standalone `reasoning_effort`, or model-specific parameter. Model IDs are never inspected or special-cased. Transport/provider failures, unavailable-model/profile failures, text/JSON contract failures, and later simulation-rule rejection remain distinct safe outcomes. The adapter never silently substitutes a model or scripted behavior. @@ -155,7 +155,7 @@ The Game API captures only schema-validated safe observations, requested world a Export requests, agent IDs, levels, ranges, and Custom dependencies are runtime-validated. Filtering and metrics remain server-owned. The export schema -is exclusively version 12; exports carrying schema version 9, 10, or 11 are +is exclusively version 13; exports carrying schema version 12 or earlier are rejected outright with no migration path. Reset clears swarm tick history and metrics while unlocking preserved roster assignments for the new experiment. diff --git a/docs/TESTING.md b/docs/TESTING.md index a3aa393..fd614b4 100644 --- a/docs/TESTING.md +++ b/docs/TESTING.md @@ -79,9 +79,9 @@ malformed-entry skipping, cache TTL, stale fallback, and safe failure states. ### Experiment archive (`packages/experiment-archive`) -`archive.test.ts` covers schema-v12-only enforcement: swarm-native provenance +`archive.test.ts` covers schema-v13-only enforcement: swarm-native provenance archival, idempotent import, query service, credential-like-data rejection before -persistence, unknown-architecture-version rejection, and non-v12 schema +persistence, unknown-architecture-version rejection, and non-v13 schema rejection. ### Game API (`apps/game-api`) diff --git a/docs/ZERO_SWARM_COMPARISON.md b/docs/ZERO_SWARM_COMPARISON.md index bd92a42..4fef9b4 100644 --- a/docs/ZERO_SWARM_COMPARISON.md +++ b/docs/ZERO_SWARM_COMPARISON.md @@ -64,7 +64,7 @@ reproducibility and exposes comparable telemetry; it does not demonstrate real-model quality, production latency, or provider cost. Real-provider studies remain explicitly opted in and should archive their safe exports separately. -The experiment archive importer now accepts only schema-v12 swarm-native +The experiment archive importer now accepts only schema-v13 swarm-native exports, which always carry swarm tick records and independent provider attempt accounting; it does not replace this same-scenario, per-tick swarm harness.