Skip to content

feat: local Laya-MLX judge backend over HTTP - #135

Closed
electroforest wants to merge 5 commits into
DevMortimer:mainfrom
electroforest:feat/laya-local-judge
Closed

electroforest wants to merge 5 commits into
DevMortimer:mainfrom
electroforest:feat/laya-local-judge

Conversation

@electroforest

Copy link
Copy Markdown

Summary

Refs #34 (local models instead of paid APIs), for the Laya-MLX case. Adds "laya" as a typesafeBackend value: every judgment runs against a local Laya-MLX model served over HTTP — no key, and judged content never leaves the machine.

Same model as #28, different transport — and #28's quality finding applies here. Ryan closed #28 after running it for hours and finding the model's answers not good. This PR does not change the model, so that concern stands. What it changes is the transport, which removes #28's operational weight:

#28 (closed) This PR
Downloads ~800 MB of weights on first use No download — talks to a server already running
Creates a venv, pip install laya-mlx, spawns a Python subprocess bridge No Python, no subprocess — plain HTTP
python/laya-bridge.py + src/laya-download.ts + Python discovery src/laya-judge.ts (~50 lines)

The endpoint speaks the same contract the TypeSafe SDK uses (POST /v1/systemone with a SystemOneRequest), so requests and answers are unchanged and every guard works through the existing judgeFor seam. Distinct from #134, which is about configurable remote Jev endpoints; this is a local model, per #34.

Design decisions

  • Consent is kept. /warden enable shows a local dialog ("nothing is sent to any server"). feat: laya-mlx local judge backend #28 bypassed consent for the local backend; this keeps it, because consent is also the master switch for judgments and the disclosure stays honest. One line to flip if you disagree.
  • A dead server behaves like any failing backend: per-request failures, then the judge cooldown, then one notice that says how to start the server. Guards fail open throughout, as with a dead remote.
  • No new Jev questions — the guard question set is untouched, so no calibration runs are needed.
  • /warden status names the local judge; getUsage() feeds the session line (requests counted, $0).

Usage

Set "typesafeBackend": "laya" in ~/.pi/agent/pi-warden/config.json, run /warden enable, and start any local server that serves POST /v1/systemone with the System One schema (state + questions in, typed answers out). Judgments then run locally at ~15–70 ms per request.

Testing

  • 6 new tests in tests/laya-judge.test.ts — wire format, model-name fallback, HTTP error status, malformed response, aborted signal, usage counter — against a stub HTTP server, so no external process is needed.
  • 2 new assertions in tests/backend.test.ts.
  • npm run check (typecheck + 1192 tests + build) is green.

Commits

Five atomic commits: judge adapter → backend registry → warden wiring → docs → version bump (0.74.0, last, per CONTRIBUTING).

Cinco added 5 commits September 28, 2026 06:31
Implements the pi-typesafe Judge contract by POSTing the same
SystemOneRequest to a local /v1/systemone endpoint — the wire shape the
TypeSafe SDK uses — so requests and answers are unchanged and every guard
works through the existing seam. No key, no model download, no Python
subprocess: the endpoint is a long-running server the user runs
themselves. A dead or failing endpoint surfaces as per-request failures
and the judge cooldown, like any failing backend.
The backend registry is the single seam every judgment flows through,
so a new backend is one accepted value: no key to resolve, no host in
the consent disclosure, and judgeOptions never forwards it to the remote
client. backendHost names it for the rules-audit and test-confirm
messages.
judgeFor returns the local judge through the same watched proxy as the
remote client, so every guard, the conscience adapters, the stuck and
done checks and the rules guard keep working unchanged. Consent is
still required — /warden enable shows a local dialog that says nothing
leaves the machine — because consent is also the master switch for
judgments. /warden status names the local judge, and a network failure
during cooldown carries a remedy that says how to start the server.
configuration.md gains the third backend value, data-handling.md says
what stays on this machine with it, and the changelog carries the
Unreleased entry.
@DevMortimer

Copy link
Copy Markdown
Owner

Thank you for this, and for taking the #28 findings seriously instead of glossing over them. Moving from a Python bridge to plain HTTP was the right instinct, and this PR pushed the project to solve the general problem.

I am closing it because that general problem is now solved upstream of pi-warden: pi-typesafe 0.8.0 accepts a caller-supplied endpoint, and pi-warden 0.74.0 lets typesafeBackend name one (#134, #136). Upstream Laya already ships a Jev-compatible server (laya-serve, POST /v1/systemone), so it works through that setting with no Laya-specific code in pi-warden. The numbers from testing it are in #134.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants