Skip to content

harness: System One, a decision loop over a typed action space (systemone, fifteenth base) - #223

Closed
richard-epsilla wants to merge 15 commits into
mainfrom
systemone-base
Closed

richard-epsilla wants to merge 15 commits into
mainfrom
systemone-base

Conversation

@richard-epsilla

Copy link
Copy Markdown
Collaborator

What

The open-source System One Harness (v0.1.1) joins as the base systemone. It drives TypeSafe's Jev, a decision model that answers typed questions with probabilities and writes no text, through OpenRouter's decisions endpoint. Not a coding CLI: a loop over a finite action space.

  • Turn process runner/systemone_driver.py. The harness's MCP server is the environment, its tools compiled to actions each turn (observe and reset are the protocol; enumerable parameters compile; a tool that needs free text is named in the trace as not offered). With no server configured, the built-in order desk. disabledTools are withheld from the question itself. The desk's state and the loop's steps persist under the workspace for continuation.
  • Relay. The route is registered at the provider's API root, because OpenRouter serves decisions at /api/alpha/decisions and not under /api/v1 (404 measured 2026-09-19). Measured through the runner's own relay: six actions, 1.44 s; the relay's served-model and usage taps agree with the driver's report (typesafe/jev-1.13-20260917, 6304 in / 1410 out).
  • A new outcome in the turn contract. A result of subtype incomplete with a reason: the loop stopped because the model asked for help, or a destructive action never cleared its confidence bar. The runner reports the status, the poll body carries the reason, the gateway puts it in incomplete_details and retries nothing.
  • Gateway. jev-1.13 and jev-latest on OpenRouter's vendor table only (added after the shared copy so TokenRouter and Vercel do not inherit them), the systemone model and base catalogs, the OpenRouter wiring.
  • Console. The base, its mark, an Instructions label. Entrypoint. A pinned venv on the data volume, proven by import and version, rebuilt when the pin moves. CI. The runner job installs the same pin so the driver tests run.

Verified

  • runner/tests/test_systemone_backend.py (9) and gateway/tests/test_systemone_catalog.py (5): route roots, placeholder never the key, the event stream, continuation, the incomplete reason, disabling by omission, the catalog rules. Whole runner and gateway suites: 1035 passed; the 17 failures are test_media_mcp.py, which needs ffmpeg on the machine that ran them.
  • One live turn through _build_systemone and the driver against OpenRouter, as above.
  • The base on the OSS test VM as a derived image from 0.20.1 on a fresh volume: see the follow-up comment.

Notes in docs/support-matrix-notes.md ("systemone") and a line in docs/harness-verification.md on why the artifact scenario does not apply.

🤖 Generated with Claude Code

… a typed action space (systemone)

The open-source System One Harness (HarnessRouter/SystemOneHarness, v0.1.1)
drives TypeSafe's Jev, a decision model that answers typed questions with
probabilities and writes no text. The base's turn process is
runner/systemone_driver.py: the harness's MCP server is the environment,
its tools compiled to actions each turn (the built-in order desk with none);
disabledTools are withheld from the question itself; the desk's state and
the loop's steps persist under the workspace for continuation. Events are
claude's stream-json, so the normaliser is the passthrough.

The route rides the loopback relay at the provider's API root, because
OpenRouter serves decisions at /api/alpha/decisions and not under /api/v1
(404 measured). One live turn: six actions, 1.44 s; the relay's served-model
and usage taps agree with the driver's own report (typesafe/jev-1.13-20260917,
6304 in / 1410 out).

New in the turn contract: a result of subtype "incomplete" with a "reason".
A loop that stops because the model asked for help, or a destructive action
never cleared its confidence bar, is neither a failure nor a step cap; the
runner reports the status, the poll body carries the reason, and the gateway
puts it in incomplete_details without retrying another connection.

Gateway: jev-1.13 and jev-latest on OpenRouter's table only, the systemone
model and base catalogs, the OpenRouter wiring. Console: the base, its mark
and an Instructions label. Entrypoint: a pinned venv on the data volume,
proven by import and version. CI installs the same pin for the driver tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
unified-harness-protocol Ready Ready Preview Sep 20, 2026 10:39am UTC

Request Review

…eaves no venv behind on failure

HR_SYSTEMONE_SPEC overrides the pip requirement (a mirror, a fork, a tree
copied into a derived image; the version proven is the pin either way). A
failed pip step used to leave the venv's python in place, and the executable
is the definition of installed, so the box reported the base available when
it could not run (fresh volume, 2026-09-19). Every failure path removes it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A follow-up that carries previous_response_id and no harness_id (the
protocol asks only for the former) was routed by the inherited model name,
and for a base whose models no chat backend serves that fell to the default
backend: a systemone follow-up asked claude for jev-1.13 (2026-09-19). The
session vertex records the harness the conversation started on; the
continuation takes it from there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

Verified on the OSS test VM, 2026-09-19

Derived image from harnessrouter/harnessrouter:0.20.1 (runner, gateway, entrypoint, the harness tree at /opt/systemone-harness with HR_SYSTEMONE_SPEC pointing at it, and the console built locally with NEXT_PUBLIC_HR_EDITION=selfhost), running as hr-s1 on a fresh volume beside hr-test, HR_BACKENDS=systemone.

First run on a fresh volume. The entrypoint installed the harness into its own venv, proved the import and the version (0.1.1), and reported the base available. The first attempt without the source override failed at git clone: the harness repository is private, so pip inside the image cannot reach the tag. That is the one open item before this ships: publish the repository or the package. The installer now removes the venv on any failure, because the first failed attempt had left the venv's python behind and the box reported the base available when it could not run.

Through the console's API, base systemone, model jev-1.13, an OpenRouter integration:

step result
"Ship order A-104 with the cheapest carrier." completed, six function_call items (pick, pick, pick, pack, choose_carrier post, ship), usage 6304 in / 1410 out, served jev-1.13, 2 s
"Pick everything for order A-104 and pack it, but do not ship yet." completed after four actions; the model then chose finish with goal-reached 0.95, shown as a reasoning item with the distribution
"Add a gift note to the order." with previous_response_id and no harness_id same session, add_note gift, then finish at 0.89: "Goal reached after 1 action."

The last row found a gateway defect and fixed it on this branch (ef42850): a continuation with no harness_id was routed by the inherited model name and fell to the claude backend. A continuation now takes the harness from its session.

In the console (screenshots at 1440 and 390): System One in the system harnesses list with its mark and jev-1.13 in the model chip; both turns render with "Used 4 tools" and "Used a tool" groups that open to Pick Item, Pick Item, Pick Item, Pack, Done and Add Note, Done, and the harness's closing sentences. Nothing overlaps or clips at phone width.

Not exercised here: a harness with an MCP server as the environment (covered by the package's tests and s1 tools on the order desk served over MCP), and the incomplete reason end to end through the console, which needs a run the model refuses; the runner and gateway tests pin the path.

…tom format can drive

Every backend's model list carried the rows of custom-endpoint integrations
as unavailable, so a System One harness's picker offered claude-sonnet-4.6,
gpt-5.5 and four more as "(no provider)" (hr-test, 2026-09-19). On a chat
backend the greyed row explains why a configured model is not pickable
there; on a backend no custom format drives it is a chat model on a harness
that cannot use one. Both the global catalog and the per-harness view now
ask _custom_can_drive first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e says so

A skill is prose an agent reads and scripts it runs from a shell, and each
built-in needs free text (an image prompt, a document's content, HTML for a
PDF). A System One model chooses among offered actions and writes nothing,
so a harness on this base showed three skills enabled that could never act
(hr-test, 2026-09-19). The base declares "skills": False in the catalog;
the bases endpoint offers none and reports takesSkills; a turn mounts none
for the built-in harness and for a fork alike; the settings page replaces
the Skills list with one sentence on where guidance and scripts go instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

Two more findings from Richard's manual pass on hr-test, both fixed on the branch and verified on the public console at 1440 and 390:

  • The picker offered chat models greyed as "(no provider)" on a System One harness (10eac7a). The models endpoint appended every custom-endpoint integration's rows to every backend as unavailable. On a chat backend that greyed row explains why a configured model is not pickable; on a backend no custom format can drive it is a chat model on a harness that cannot use one. Both the catalog and the per-harness view now ask _custom_can_drive first. A System One harness lists jev-1.13 and jev-latest and nothing else.
  • Three built-in skills showed enabled on a System One harness (40697ae). A skill is prose an agent reads and scripts it runs from a shell, and each built-in needs free text; a System One model chooses among offered actions and writes nothing. The base declares "skills": False; the bases endpoint offers none and reports takesSkills; a turn mounts none for the built-in harness and a fork alike; the settings page replaces the list with one sentence on where guidance and scripts go instead.

hr-test runs the derived image with both (harnessrouter:systemone-40697ae-ui).

richard-epsilla and others added 2 commits September 19, 2026 19:53
0.2.0 adds the browser environment on Browser Use and moves to the 2.x MCP
SDK. The base's venv is its own on the data volume, so the SDK major there
is independent of the runner's pin.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The venv installs the browser extra (Browser Use) and a Playwright
Chromium beside it on the data volume, world-readable, with
PLAYWRIGHT_BROWSERS_PATH exported for every session process: the image
ships no browser and a session uid cannot read root's cache. That is what
the browser environment and the Super Mario starter kit's environment run
on. Pins move to 0.2.1 (meta.risk for tools, canvas mode).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

Two more commits on this branch, driven by the fifth starter kit (HarnessRouter/starter-kit#27, Super Mario on the System One base):

  • 5fed672, the base gets a browser. The venv installs the harness's browser extra (Browser Use) and a Playwright Chromium beside it on the data volume, world-readable, with PLAYWRIGHT_BROWSERS_PATH exported for every session process: the image ships no browser and a session uid cannot read root's cache. Pins move to harness 0.2.1 (a tool may declare its risk in its MCP meta; canvas mode on the browser environment).
  • The kit's environment is a plugin MCP server (stdio, ./bin/mario-env) that runs on that venv and browser; the runner already launches it per turn and the driver takes it as the environment. Verified on hr-test from a launch on the public console: base systemone, plugin mario-env@0.1.0 with its server and nothing skipped, frames flowing to the kit's page, about one action a second.

hr-test runs harnessrouter:systemone-5fed672-kit, the branch plus the kit baked in as install-kits.sh would.

…ner, not once per volume

The System One base's Chromium lives on the data volume, but the libraries it links against were
installed by `playwright install --with-deps` into the container that did the first run and
nowhere else. A container recreated over that volume (a new image version, a redeploy) had the
binary and none of its libraries: libatk, libatspi and libXcomposite "not found", every launch
dead before CDP came up, and every run of the Super Mario kit failing at reset. Measured on the
test box after a redeploy.

The libraries are now their own step, keyed on a marker in the container's own filesystem: taken
on the first install, and on every start that finds the venv already there. A failure on a later
start is a warning that names the consequence and is retried on the next start; on the first
install it fails the install, as a browser that cannot start should.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

One more entrypoint change, cdcd738: Chromium's system libraries are installed per container.

The browser lives on the data volume, but playwright install --with-deps put the libraries it links against into the container that did the first run and nowhere else. A container recreated over that volume (a new image version, a redeploy) had the binary and none of its libraries: libatk, libatspi and libXcomposite "not found", every launch dead before CDP came up, every Super Mario run failing at reset. Measured on the test box after a redeploy.

Now the libraries are their own step keyed on a marker in the container's own filesystem, taken on the first install and on every start that finds the venv already there (about 12 s on a recreated container). Verified both ways: a redeploy over the existing volume, and a brand-new volume installing the harness from the repository, which is public now, so the default spec resolves without an override.

Its actions are its environment's: the tools of the MCP server the harness configures, compiled
at the start of every turn. Listing the order desk's six actions in the base catalog showed them
on the Super Mario harness as "built into System One", enabled, while that harness never offers
them (its environment is the kit's plugin). The order desk stays what the loop runs without a
server, as a demonstration; a tool disabled on a harness is still withheld by omission from the
question the model answers, whatever environment offers it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

68be604: the System One base declares no built-in tools. Its actions are the environment's (the MCP server the harness configures); listing the order desk's six showed them on the Super Mario harness as built into System One, enabled, while that harness never offers them. The settings page now reads 0 configured tools for it; the order desk stays what the loop runs without a server.

A kit knows how long its turns run: a game played several decisions a second needs more steps
than a document does. kit.json harness.max_step and harness.timeout_seconds now reach the
HarnessBody a launch builds; absent, the base's defaults stand.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The extra wants mcp 2.1 and the runner pins 1.28.1; the tests use the recorded
provider and the built-in order environment, and the MCP client is imported only
when an MCP environment is opened. In the image the harness has its own
environment on the volume, so nothing there changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The System One Harness now has a logo of its own (assets/logo.png in its
repository); the placeholder SVG goes, and the console's base catalog shows
the mark at 512 by 512, the size the largest existing base logo ships at.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@richard-epsilla

Copy link
Copy Markdown
Collaborator Author

Closing in favour of a bot-authored successor on the same branch: the main ruleset needs an approving review from a writer other than the author, and richard-epsilla is both the author here and the only writer. Nothing about the branch changes; the successor carries the same commits plus the harness's own logo (9037758).

This branch was successfully deployed

1 active deployment
Preview — 90377588 Deployed Sep 20, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant