laya-shim runs a Laya
checkpoint behind TypeSafe's System One route, POST /v1/systemone. That
route is the one omp calls for its judge model role. Point the role at this
server and omp's typed yes/no, choice, and score decisions run on your machine
instead of on TypeSafe's Jev.
The shim runs one of two backends:
LAYA_BACKEND |
Library | Runs on |
|---|---|---|
mlx (default) |
Laya-MLX, an MLX port | Apple silicon GPU |
torch |
Laya, the upstream PyTorch release | Any platform PyTorch supports |
Both libraries load the same Hugging Face checkpoints, and their predict()
functions take and return the same JSON. You can switch backends without
changing omp's config.
Upstream Laya also ships its own Jev-compatible server, laya-serve. This shim
exists so that both backends run behind one server with one set of settings.
Install mise. mise installs uv, and uv installs Python and the backend's packages the first time you start the server.
mise run serveOn the first run, the server downloads the checkpoint from Hugging Face. The typed-decisions checkpoint is about 850 MB. Later runs load it from the local cache. On an Apple silicon Mac, that took about 1 second with MLX and 18 seconds with PyTorch. The server is ready when it logs this line:
INFO laya-shim: listening on http://127.0.0.1:8765/v1/systemone
To check it by hand, send a request:
curl -s localhost:8765/v1/systemone -d '{
"state": "All tests pass now.",
"questions": {"claims": {"type": "noul", "instructions": "Does the reply claim tests pass?"}}
}'To stop the server, press Ctrl+C or send it SIGTERM. The server finishes the
request it's answering, closes its socket, and logs stopped. After Ctrl+C,
mise run exits with status 130, the usual status for an interrupted command.
If you press Ctrl+C during startup, for example during a download, the server
logs interrupted during startup and exits.
ui/ holds a test page that sends System One requests to a running server.
Start the server, and then run this command in a second terminal:
mise run uiOpen http://localhost:3000. Pick a sample from the list, or edit the state and the questions JSON to write your own, and then click Ask. The page shows each answer, its confidence, and its probabilities, and it keeps the raw response under Raw response.
Each sample is a request that an omp feature sends, loaded from
tests/fixtures. omp sends the state as text, a
JSON object, or a JSON array. The page does the same: if the state box holds a
JSON object or array, the page sends it as JSON, and otherwise it sends the
text.
The page's Bun server forwards /api/systemone to the shim at LAYA_HOST and
LAYA_PORT, because the shim sends no CORS headers. To serve the page on
another port, set UI_PORT.
mise.toml sets these variables. The shim uses the same defaults when you run
it without mise.
| Variable | Default | Meaning |
|---|---|---|
LAYA_BACKEND |
mlx |
mlx or torch. |
LAYA_MODEL |
convaiinnovations/laya |
Hugging Face repository or local path. |
LAYA_SUBFOLDER |
typed-decisions |
Checkpoint inside LAYA_MODEL. Leave it empty for English, or set multilingual or typed-decisions. |
LAYA_HOST |
127.0.0.1 |
Bind address. |
LAYA_PORT |
8765 |
Bind port. |
LAYA_LOG_COLOR |
auto |
auto colors log levels when standard error is a terminal and NO_COLOR isn't set. Set always or never to override that. |
To change a value for this checkout only, create mise.local.toml. Git ignores
this file.
[env]
LAYA_BACKEND = "torch"To change a value for one run, set it in your shell:
LAYA_BACKEND=torch mise run serveA value from your shell or from mise.local.toml takes precedence over the
default in mise.toml.
When you switch backends, uv removes the other backend's packages and installs the ones you selected. Both backends read the same Hugging Face cache, so switching doesn't download the checkpoint again.
LAYA_SUBFOLDER |
Context | Notes |
|---|---|---|
| (empty) | 512 tokens | English only. |
multilingual |
1,024 tokens | More than 100 languages. |
typed-decisions |
1,024 tokens | Fine-tuned on typed decisions. On upstream's typed-decisions benchmark it scores 0.766, against 0.362 for the English checkpoint. |
-
Add a provider to
~/.omp/agent/models.yml:providers: laya: baseUrl: http://127.0.0.1:8765 api: typesafe apiKey: laya-local models: - id: laya-local name: Laya (local)
api: typesafemakes omp send judge requests to{baseUrl}/v1/systemone. omp skips a judge that has no API key, soapiKeyneeds a value. The shim ignores it. -
Select the model for the
judgerole in~/.omp/agent/config.yml:modelRoles: judge: laya/laya-local retry: fallbackChains: judge: - typesafe/jev-latest
If the shim is down or rejects a request, omp tries the next judge in
fallbackChains.judge. To keep every judgment on your machine, setjudge: []instead. -
Start the server, and then run a command that uses the judge:
omp find "where the shim returns HTTP 422" .
Each judgment adds an access log line to the server's output:
INFO laya-shim: 127.0.0.1 "POST /v1/systemone" 200 questions=3 352.5 ms
omp treats any model with api: typesafe as a native System One judge, the
same as Jev. Two settings that default to auto turn on because of this:
find.enabledadds thefindtool.ttsr.judgeasks judged rulebook rules about each finished reply and tool call.
These limits come from Laya, not from the shim. Check them before you rely on the answers.
- Short context. Jev reads about 33,000 tokens. Laya reads 512 or 1,024,
and that budget covers the question, the options, and the state. Laya cuts
the end off a longer state without an error. omp sends long states for
judged rulebook rules, git staging, and
find, so in those cases Laya only sees the start. If judged rules give bad answers, setttsr.judge: off. - Few choice options. All options for one question share a budget of 192
to 256 tokens. A question with too many options fails, and the shim returns
422. omp doesn't retry a422and moves on to the next judge in the chain. - Weak yes/no answers. Upstream reports that
noulquestions can follow the option labels instead of the input, most often on the English checkpoint. In a test with the typed-decisions checkpoint, "All tests pass now." scored 0.54 for "Does the reply claim tests pass?", which is close to a coin toss. - Latency grows with input length. A full 1,024-token state takes about
13 times as long as a short one. omp's
findsends whole files, so expect hundreds of milliseconds per request there. See Benchmarks.
benchmarks/bench.py sends System One requests to a running server and times
each HTTP round trip. To measure your machine, start the server, and then run
this command in a second terminal:
mise run benchEach workload runs 5 untimed warmup requests and then 50 timed ones. To change the counts, run this command:
uv run --no-project python benchmarks/bench.py --iterations 100 --warmup 10The long workload's state is about 12,700 characters, far more than Laya's 1,024-token context holds, so Laya reads a full context and drops the rest.
These results come from an Apple M2 Pro with 16 GB of memory and the typed-decisions checkpoint, measured on September 25, 2026. Both backends ran on the GPU: MLX through Metal, and PyTorch through MPS.
| Workload | MLX p50 | MLX p95 | PyTorch p50 | PyTorch p95 |
|---|---|---|---|---|
| Short state, 1 question | 17.2 ms | 19.0 ms | 34.0 ms | 37.3 ms |
| Short state, 3 questions | 42.0 ms | 44.3 ms | 56.9 ms | 60.3 ms |
| Short state, 10 questions | 91.5 ms | 94.3 ms | 139.8 ms | 150.6 ms |
| Long state, 1 question | 232.1 ms | 248.2 ms | 301.8 ms | 324.1 ms |
MLX was about twice as fast on a single short question and 1.3 to 1.5 times as fast on the rest. On load time, MLX took about 1 second, and PyTorch took 8 to 18 seconds.
The server writes these log lines to standard error:
- Startup settings, and when the checkpoint starts and finishes loading.
- One
httpxline for each GET request to Hugging Face, which includes file downloads. The Hugging Face client also draws a progress bar for each file. - Warnings from Laya, one line each, such as the calibration temperature it clamps at load time.
- One access log line for each request, with the status, question count, and time.
- The reason for each rejected request, and a traceback for each failed inference.
When color is on, INFO is green, WARNING is yellow, and ERROR is red.
The project uses Python's standard src layout:
src/laya_shim/server.pyis the server.pyproject.tomlmaps thelaya-shimcommand to itsmain()function.benchmarks/bench.pyis the latency benchmark. It isn't part of the package.ui/is the browser test page. It needs Bun, which mise installs, and has no other dependencies.tests/holds the tests and the omp request fixtures.
To run the server without mise, run uv run --extra mlx laya-shim.
hk runs ruff check and then ruff format as a
pre-commit hook. mise installs the hook when a shell with mise activate
enters this directory. To install it by hand, run hk install --mise. The hook
fixes what it can and stages the result. It blocks the commit if an error
remains that ruff can't fix. To skip it for one commit, run
HK=0 git commit.
To run the same checks on every file, run hk check --all. To apply the fixes,
run hk fix --all.
The CI workflow in .github/workflows/ci.yml runs hk check --all,
uv lock --check, and the tests that don't need the checkpoint, on each push
to master and on each pull request. CI skips the contract tests, because
they need an 850 MB checkpoint download.
Every action in the workflows is pinned to a full commit SHA, with the version
in a trailing comment. Dependabot updates those pins and the uv dependencies
every week. Dependabot doesn't read mise.toml, so update the versions of hk,
ruff, and uv there yourself.
To run every test, run this command:
mise run testThe tests start their own server on a free port, so you don't need
mise run serve running. They use the backend that LAYA_BACKEND selects.
tests/test_contract.pyloads the checkpoint and sends each request intests/fixtures. Each fixture copies a request from the omp feature that itssourcefield names, such as auto thinking, judged rules, git AI staging,find, and the evaljudge()helper. The test checks each reply against omp's System One types: every question has an answer of the same type, probabilities cover every option or level and sum to 1, a choice is the most likely option, and a score is the weighted mean of its levels.tests/test_http.pychecks the status codes omp acts on, with a stub in place of Laya. omp retries a5xxresponse, and it moves on to the next judge after a4xx.
To run only the tests that don't need the checkpoint, run
uv run pytest -m "not model".
When omp changes a request, update the matching fixture. The test page loads the same files, so it picks up the change too.
Copyright 2026 Robbie Blaine. This project is licensed under the Apache
License, Version 2.0. For the full text, see LICENSE.
Laya, Laya-MLX, and the Laya checkpoints are separate projects with their own licenses, all of which are also Apache-2.0.