harness: System One, a decision loop over a typed action space (systemone, fifteenth base) - #223
richard-epsilla wants to merge 15 commits into
Conversation
… a typed action space (systemone) The open-source System One Harness (HarnessRouter/SystemOneHarness, v0.1.1) drives TypeSafe's Jev, a decision model that answers typed questions with probabilities and writes no text. The base's turn process is runner/systemone_driver.py: the harness's MCP server is the environment, its tools compiled to actions each turn (the built-in order desk with none); disabledTools are withheld from the question itself; the desk's state and the loop's steps persist under the workspace for continuation. Events are claude's stream-json, so the normaliser is the passthrough. The route rides the loopback relay at the provider's API root, because OpenRouter serves decisions at /api/alpha/decisions and not under /api/v1 (404 measured). One live turn: six actions, 1.44 s; the relay's served-model and usage taps agree with the driver's own report (typesafe/jev-1.13-20260917, 6304 in / 1410 out). New in the turn contract: a result of subtype "incomplete" with a "reason". A loop that stops because the model asked for help, or a destructive action never cleared its confidence bar, is neither a failure nor a step cap; the runner reports the status, the poll body carries the reason, and the gateway puts it in incomplete_details without retrying another connection. Gateway: jev-1.13 and jev-latest on OpenRouter's table only, the systemone model and base catalogs, the OpenRouter wiring. Console: the base, its mark and an Instructions label. Entrypoint: a pinned venv on the data volume, proven by import and version. CI installs the same pin for the driver tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
…eaves no venv behind on failure HR_SYSTEMONE_SPEC overrides the pip requirement (a mirror, a fork, a tree copied into a derived image; the version proven is the pin either way). A failed pip step used to leave the venv's python in place, and the executable is the definition of installed, so the box reported the base available when it could not run (fresh volume, 2026-09-19). Every failure path removes it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A follow-up that carries previous_response_id and no harness_id (the protocol asks only for the former) was routed by the inherited model name, and for a base whose models no chat backend serves that fell to the default backend: a systemone follow-up asked claude for jev-1.13 (2026-09-19). The session vertex records the harness the conversation started on; the continuation takes it from there. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Verified on the OSS test VM, 2026-09-19Derived image from First run on a fresh volume. The entrypoint installed the harness into its own venv, proved the import and the version (0.1.1), and reported the base available. The first attempt without the source override failed at Through the console's API, base
The last row found a gateway defect and fixed it on this branch (ef42850): a continuation with no In the console (screenshots at 1440 and 390): System One in the system harnesses list with its mark and Not exercised here: a harness with an MCP server as the environment (covered by the package's tests and |
…tom format can drive Every backend's model list carried the rows of custom-endpoint integrations as unavailable, so a System One harness's picker offered claude-sonnet-4.6, gpt-5.5 and four more as "(no provider)" (hr-test, 2026-09-19). On a chat backend the greyed row explains why a configured model is not pickable there; on a backend no custom format drives it is a chat model on a harness that cannot use one. Both the global catalog and the per-harness view now ask _custom_can_drive first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e says so A skill is prose an agent reads and scripts it runs from a shell, and each built-in needs free text (an image prompt, a document's content, HTML for a PDF). A System One model chooses among offered actions and writes nothing, so a harness on this base showed three skills enabled that could never act (hr-test, 2026-09-19). The base declares "skills": False in the catalog; the bases endpoint offers none and reports takesSkills; a turn mounts none for the built-in harness and for a fork alike; the settings page replaces the Skills list with one sentence on where guidance and scripts go instead. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Two more findings from Richard's manual pass on hr-test, both fixed on the branch and verified on the public console at 1440 and 390:
hr-test runs the derived image with both ( |
0.2.0 adds the browser environment on Browser Use and moves to the 2.x MCP SDK. The base's venv is its own on the data volume, so the SDK major there is independent of the runner's pin. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The venv installs the browser extra (Browser Use) and a Playwright Chromium beside it on the data volume, world-readable, with PLAYWRIGHT_BROWSERS_PATH exported for every session process: the image ships no browser and a session uid cannot read root's cache. That is what the browser environment and the Super Mario starter kit's environment run on. Pins move to 0.2.1 (meta.risk for tools, canvas mode). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Two more commits on this branch, driven by the fifth starter kit (HarnessRouter/starter-kit#27, Super Mario on the System One base):
hr-test runs |
…ner, not once per volume The System One base's Chromium lives on the data volume, but the libraries it links against were installed by `playwright install --with-deps` into the container that did the first run and nowhere else. A container recreated over that volume (a new image version, a redeploy) had the binary and none of its libraries: libatk, libatspi and libXcomposite "not found", every launch dead before CDP came up, and every run of the Super Mario kit failing at reset. Measured on the test box after a redeploy. The libraries are now their own step, keyed on a marker in the container's own filesystem: taken on the first install, and on every start that finds the venv already there. A failure on a later start is a warning that names the consequence and is retried on the next start; on the first install it fails the install, as a browser that cannot start should. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
One more entrypoint change, cdcd738: Chromium's system libraries are installed per container. The browser lives on the data volume, but Now the libraries are their own step keyed on a marker in the container's own filesystem, taken on the first install and on every start that finds the venv already there (about 12 s on a recreated container). Verified both ways: a redeploy over the existing volume, and a brand-new volume installing the harness from the repository, which is public now, so the default spec resolves without an override. |
Its actions are its environment's: the tools of the MCP server the harness configures, compiled at the start of every turn. Listing the order desk's six actions in the base catalog showed them on the Super Mario harness as "built into System One", enabled, while that harness never offers them (its environment is the kit's plugin). The order desk stays what the loop runs without a server, as a demonstration; a tool disabled on a harness is still withheld by omission from the question the model answers, whatever environment offers it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
68be604: the System One base declares no built-in tools. Its actions are the environment's (the MCP server the harness configures); listing the order desk's six showed them on the Super Mario harness as built into System One, enabled, while that harness never offers them. The settings page now reads 0 configured tools for it; the order desk stays what the loop runs without a server. |
A kit knows how long its turns run: a game played several decisions a second needs more steps than a document does. kit.json harness.max_step and harness.timeout_seconds now reach the HarnessBody a launch builds; absent, the base's defaults stand. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The extra wants mcp 2.1 and the runner pins 1.28.1; the tests use the recorded provider and the built-in order environment, and the MCP client is imported only when an MCP environment is opened. In the image the harness has its own environment on the volume, so nothing there changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The System One Harness now has a logo of its own (assets/logo.png in its repository); the placeholder SVG goes, and the console's base catalog shows the mark at 512 by 512, the size the largest existing base logo ships at. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Closing in favour of a bot-authored successor on the same branch: the main ruleset needs an approving review from a writer other than the author, and richard-epsilla is both the author here and the only writer. Nothing about the branch changes; the successor carries the same commits plus the harness's own logo (9037758). |
What
The open-source System One Harness (v0.1.1) joins as the base
systemone. It drives TypeSafe's Jev, a decision model that answers typed questions with probabilities and writes no text, through OpenRouter's decisions endpoint. Not a coding CLI: a loop over a finite action space.runner/systemone_driver.py. The harness's MCP server is the environment, its tools compiled to actions each turn (observe and reset are the protocol; enumerable parameters compile; a tool that needs free text is named in the trace as not offered). With no server configured, the built-in order desk.disabledToolsare withheld from the question itself. The desk's state and the loop's steps persist under the workspace for continuation./api/alpha/decisionsand not under/api/v1(404 measured 2026-09-19). Measured through the runner's own relay: six actions, 1.44 s; the relay's served-model and usage taps agree with the driver's report (typesafe/jev-1.13-20260917, 6304 in / 1410 out).incompletewith areason: the loop stopped because the model asked for help, or a destructive action never cleared its confidence bar. The runner reports the status, the poll body carries the reason, the gateway puts it inincomplete_detailsand retries nothing.jev-1.13andjev-lateston OpenRouter's vendor table only (added after the shared copy so TokenRouter and Vercel do not inherit them), the systemone model and base catalogs, the OpenRouter wiring.Verified
runner/tests/test_systemone_backend.py(9) andgateway/tests/test_systemone_catalog.py(5): route roots, placeholder never the key, the event stream, continuation, the incomplete reason, disabling by omission, the catalog rules. Whole runner and gateway suites: 1035 passed; the 17 failures aretest_media_mcp.py, which needs ffmpeg on the machine that ran them._build_systemoneand the driver against OpenRouter, as above.Notes in
docs/support-matrix-notes.md("systemone") and a line indocs/harness-verification.mdon why the artifact scenario does not apply.🤖 Generated with Claude Code