Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

laya-shim

laya-shim runs a Laya checkpoint behind TypeSafe's System One route, POST /v1/systemone. That route is the one omp calls for its judge model role. Point the role at this server and omp's typed yes/no, choice, and score decisions run on your machine instead of on TypeSafe's Jev.

The shim runs one of two backends:

LAYA_BACKEND Library Runs on
mlx (default) Laya-MLX, an MLX port Apple silicon GPU
torch Laya, the upstream PyTorch release Any platform PyTorch supports

Both libraries load the same Hugging Face checkpoints, and their predict() functions take and return the same JSON. You can switch backends without changing omp's config.

Upstream Laya also ships its own Jev-compatible server, laya-serve. This shim exists so that both backends run behind one server with one set of settings.

Before you begin

Install mise. mise installs uv, and uv installs Python and the backend's packages the first time you start the server.

Start the server

mise run serve

On the first run, the server downloads the checkpoint from Hugging Face. The typed-decisions checkpoint is about 850 MB. Later runs load it from the local cache. On an Apple silicon Mac, that took about 1 second with MLX and 18 seconds with PyTorch. The server is ready when it logs this line:

INFO laya-shim: listening on http://127.0.0.1:8765/v1/systemone

To check it by hand, send a request:

curl -s localhost:8765/v1/systemone -d '{
  "state": "All tests pass now.",
  "questions": {"claims": {"type": "noul", "instructions": "Does the reply claim tests pass?"}}
}'

To stop the server, press Ctrl+C or send it SIGTERM. The server finishes the request it's answering, closes its socket, and logs stopped. After Ctrl+C, mise run exits with status 130, the usual status for an interrupted command. If you press Ctrl+C during startup, for example during a download, the server logs interrupted during startup and exits.

Try questions in a browser

ui/ holds a test page that sends System One requests to a running server. Start the server, and then run this command in a second terminal:

mise run ui

Open http://localhost:3000. Pick a sample from the list, or edit the state and the questions JSON to write your own, and then click Ask. The page shows each answer, its confidence, and its probabilities, and it keeps the raw response under Raw response.

Each sample is a request that an omp feature sends, loaded from tests/fixtures. omp sends the state as text, a JSON object, or a JSON array. The page does the same: if the state box holds a JSON object or array, the page sends it as JSON, and otherwise it sends the text.

The page's Bun server forwards /api/systemone to the shim at LAYA_HOST and LAYA_PORT, because the shim sends no CORS headers. To serve the page on another port, set UI_PORT.

Configure the server

mise.toml sets these variables. The shim uses the same defaults when you run it without mise.

Variable Default Meaning
LAYA_BACKEND mlx mlx or torch.
LAYA_MODEL convaiinnovations/laya Hugging Face repository or local path.
LAYA_SUBFOLDER typed-decisions Checkpoint inside LAYA_MODEL. Leave it empty for English, or set multilingual or typed-decisions.
LAYA_HOST 127.0.0.1 Bind address.
LAYA_PORT 8765 Bind port.
LAYA_LOG_COLOR auto auto colors log levels when standard error is a terminal and NO_COLOR isn't set. Set always or never to override that.

To change a value for this checkout only, create mise.local.toml. Git ignores this file.

[env]
LAYA_BACKEND = "torch"

To change a value for one run, set it in your shell:

LAYA_BACKEND=torch mise run serve

A value from your shell or from mise.local.toml takes precedence over the default in mise.toml.

When you switch backends, uv removes the other backend's packages and installs the ones you selected. Both backends read the same Hugging Face cache, so switching doesn't download the checkpoint again.

Choose a checkpoint

LAYA_SUBFOLDER Context Notes
(empty) 512 tokens English only.
multilingual 1,024 tokens More than 100 languages.
typed-decisions 1,024 tokens Fine-tuned on typed decisions. On upstream's typed-decisions benchmark it scores 0.766, against 0.362 for the English checkpoint.

Connect omp

  1. Add a provider to ~/.omp/agent/models.yml:

    providers:
      laya:
        baseUrl: http://127.0.0.1:8765
        api: typesafe
        apiKey: laya-local
        models:
          - id: laya-local
            name: Laya (local)

    api: typesafe makes omp send judge requests to {baseUrl}/v1/systemone. omp skips a judge that has no API key, so apiKey needs a value. The shim ignores it.

  2. Select the model for the judge role in ~/.omp/agent/config.yml:

    modelRoles:
      judge: laya/laya-local
    retry:
      fallbackChains:
        judge:
          - typesafe/jev-latest

    If the shim is down or rejects a request, omp tries the next judge in fallbackChains.judge. To keep every judgment on your machine, set judge: [] instead.

  3. Start the server, and then run a command that uses the judge:

    omp find "where the shim returns HTTP 422" .

    Each judgment adds an access log line to the server's output:

    INFO laya-shim: 127.0.0.1 "POST /v1/systemone" 200 questions=3 352.5 ms
    

omp treats any model with api: typesafe as a native System One judge, the same as Jev. Two settings that default to auto turn on because of this:

  • find.enabled adds the find tool.
  • ttsr.judge asks judged rulebook rules about each finished reply and tool call.

Limits

These limits come from Laya, not from the shim. Check them before you rely on the answers.

  • Short context. Jev reads about 33,000 tokens. Laya reads 512 or 1,024, and that budget covers the question, the options, and the state. Laya cuts the end off a longer state without an error. omp sends long states for judged rulebook rules, git staging, and find, so in those cases Laya only sees the start. If judged rules give bad answers, set ttsr.judge: off.
  • Few choice options. All options for one question share a budget of 192 to 256 tokens. A question with too many options fails, and the shim returns 422. omp doesn't retry a 422 and moves on to the next judge in the chain.
  • Weak yes/no answers. Upstream reports that noul questions can follow the option labels instead of the input, most often on the English checkpoint. In a test with the typed-decisions checkpoint, "All tests pass now." scored 0.54 for "Does the reply claim tests pass?", which is close to a coin toss.
  • Latency grows with input length. A full 1,024-token state takes about 13 times as long as a short one. omp's find sends whole files, so expect hundreds of milliseconds per request there. See Benchmarks.

Benchmarks

benchmarks/bench.py sends System One requests to a running server and times each HTTP round trip. To measure your machine, start the server, and then run this command in a second terminal:

mise run bench

Each workload runs 5 untimed warmup requests and then 50 timed ones. To change the counts, run this command:

uv run --no-project python benchmarks/bench.py --iterations 100 --warmup 10

The long workload's state is about 12,700 characters, far more than Laya's 1,024-token context holds, so Laya reads a full context and drops the rest.

These results come from an Apple M2 Pro with 16 GB of memory and the typed-decisions checkpoint, measured on September 25, 2026. Both backends ran on the GPU: MLX through Metal, and PyTorch through MPS.

Workload MLX p50 MLX p95 PyTorch p50 PyTorch p95
Short state, 1 question 17.2 ms 19.0 ms 34.0 ms 37.3 ms
Short state, 3 questions 42.0 ms 44.3 ms 56.9 ms 60.3 ms
Short state, 10 questions 91.5 ms 94.3 ms 139.8 ms 150.6 ms
Long state, 1 question 232.1 ms 248.2 ms 301.8 ms 324.1 ms

MLX was about twice as fast on a single short question and 1.3 to 1.5 times as fast on the rest. On load time, MLX took about 1 second, and PyTorch took 8 to 18 seconds.

Logs

The server writes these log lines to standard error:

  • Startup settings, and when the checkpoint starts and finishes loading.
  • One httpx line for each GET request to Hugging Face, which includes file downloads. The Hugging Face client also draws a progress bar for each file.
  • Warnings from Laya, one line each, such as the calibration temperature it clamps at load time.
  • One access log line for each request, with the status, question count, and time.
  • The reason for each rejected request, and a traceback for each failed inference.

When color is on, INFO is green, WARNING is yellow, and ERROR is red.

Development

The project uses Python's standard src layout:

  • src/laya_shim/server.py is the server. pyproject.toml maps the laya-shim command to its main() function.
  • benchmarks/bench.py is the latency benchmark. It isn't part of the package.
  • ui/ is the browser test page. It needs Bun, which mise installs, and has no other dependencies.
  • tests/ holds the tests and the omp request fixtures.

To run the server without mise, run uv run --extra mlx laya-shim.

hk runs ruff check and then ruff format as a pre-commit hook. mise installs the hook when a shell with mise activate enters this directory. To install it by hand, run hk install --mise. The hook fixes what it can and stages the result. It blocks the commit if an error remains that ruff can't fix. To skip it for one commit, run HK=0 git commit.

To run the same checks on every file, run hk check --all. To apply the fixes, run hk fix --all.

The CI workflow in .github/workflows/ci.yml runs hk check --all, uv lock --check, and the tests that don't need the checkpoint, on each push to master and on each pull request. CI skips the contract tests, because they need an 850 MB checkpoint download.

Every action in the workflows is pinned to a full commit SHA, with the version in a trailing comment. Dependabot updates those pins and the uv dependencies every week. Dependabot doesn't read mise.toml, so update the versions of hk, ruff, and uv there yourself.

Test omp compatibility

To run every test, run this command:

mise run test

The tests start their own server on a free port, so you don't need mise run serve running. They use the backend that LAYA_BACKEND selects.

  • tests/test_contract.py loads the checkpoint and sends each request in tests/fixtures. Each fixture copies a request from the omp feature that its source field names, such as auto thinking, judged rules, git AI staging, find, and the eval judge() helper. The test checks each reply against omp's System One types: every question has an answer of the same type, probabilities cover every option or level and sum to 1, a choice is the most likely option, and a score is the weighted mean of its levels.
  • tests/test_http.py checks the status codes omp acts on, with a stub in place of Laya. omp retries a 5xx response, and it moves on to the next judge after a 4xx.

To run only the tests that don't need the checkpoint, run uv run pytest -m "not model".

When omp changes a request, update the matching fixture. The test page loads the same files, so it picks up the change too.

License

Copyright 2026 Robbie Blaine. This project is licensed under the Apache License, Version 2.0. For the full text, see LICENSE.

Laya, Laya-MLX, and the Laya checkpoints are separate projects with their own licenses, all of which are also Apache-2.0.

About

A thin shim for Laya/Laya-MLX

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages