| title | SWAY | |
|---|---|---|
| emoji | 🎯 | |
| colorFrom | blue | |
| colorTo | gray | |
| sdk | gradio | |
| sdk_version | 6.28.0 | |
| python_version | 3.12 | |
| app_file | app.py | |
| pinned | false | |
| license | mit | |
| short_description | A word game against the Laya decision engine | |
| models |
|
A word game played against Laya, an open non-autoregressive decision engine. Each case sets a goal on the engine's answers. You write the message, the engine judges it in one forward pass, and you see every probability it produced.
The cases are real production workflows: support triage, content moderation, LLM guardrails, model routing and email security. Playing them is a hands-on way to see what typed decisions, calibrated confidence and language routing do.
| # | Case | Goal |
|---|---|---|
| 1 | First contact | Get routed to billing without billing words |
| 2 | Cold anger | Sound clearly annoyed while staying non-toxic |
| 3 | The quiet exit | Trigger churn risk while sounding calm |
| 4 | Tick tock | Convey urgency with no alarm words |
| 5 | Babel desk | Report an outage in a language other than English |
| 6 | Fog machine | Make billing and technical tie |
| 7 | Gatekeeper | Ask a real security question the guardrail does not flag |
| 8 | Short and hard | Get rated hard in 70 characters or fewer |
| 9 | Real, not phish | Write an urgent IT email that is not flagged as phishing |
| 10 | Grand finale | A calm non-English refund request with no churn and no toxicity |
Five shots per case. A shot blocked by a rule (banned word, length, language) is free. Scores reward passing with margin; three stars need a strong pass on the first shot. Every case carries a hint; using it caps that case at two stars. Progress is kept in the browser.
Each case also names the product it comes from, so the case card says where that exact decision runs.
The Lab tab runs any preset or custom question set on any input and shows the full answer.
The Autoplay tab plays the game by itself, with Laya making every decision. Pick one case or all of them, press start, and watch the search stream.
For each case the player composes eight candidate messages from a library of phrase fragments, runs them through the same sanitiser and rule checks a human shot goes through, and drops the blocked ones for free. The survivors are scored by the engine. It keeps the best three and mutates them: swap a fragment, drop a clause, add a closing line. It stops at the first candidate scoring 95 or more, or when the case runs out of model calls.
| Setting | Meaning |
|---|---|
| Quick | 12 model calls per case |
| Thorough | 30 model calls per case |
| Seed | Same seed, same search. Default 42 |
A whole run stops at 300 model calls or six minutes, whichever comes first. The final card reports cases solved, average best score, model calls, average decision time and total time, and compares the machine's best scores to yours. Your own progress is only read, never changed.
Non-English cases get French and Spanish fragments, so a candidate is never a mix of two languages and the router always has a full sentence to detect.
| Use case | What the engine decides |
|---|---|
| Support triage | Which queue a ticket belongs to, and how frustrated the customer sounds |
| Moderation | Whether a post is toxic, in the same pass that reads it for anything else |
| LLM guardrails | Whether a prompt is a jailbreak or an injection, before the model call |
| Model routing | How hard a request is, so easy ones go to a small model |
| Email security | Whether an urgent mail is genuine or phishing |
| Multilingual routing | Which checkpoint answers, from script and language detected first |
| Confidence-gated automation | Act above a threshold, send everything under it to a person |
python -m venv .venv && source .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cpu # or your CUDA build
pip install -r requirements-dev.txt
python app.py # http://localhost:7860The first start downloads the checkpoints (about 1.5 GB for English plus multilingual).
| Variable | Default | Meaning |
|---|---|---|
LAYA_CHECKPOINTS |
english,multilingual |
Checkpoints to load. multilingual alone runs in about 2.5 GB RAM |
LAYA_DEVICE |
auto | cpu, cuda or mps |
LAYA_LOW_MEMORY |
1 |
Build models on the meta device to cut peak RAM during load |
SWAY_CONCURRENCY |
2 |
Parallel model calls |
SWAY_SKIP_ENGINE |
unset | 1 starts the UI without loading models (CI) |
pytest -q # fast, no model
SWAY_MODEL_TESTS=1 pytest -q tests/test_levels_model.py # every case has a known passing answer,
# and Quick autoplay solves at least 6 of 10The Lab is exposed as /decide:
from gradio_client import Client
c = Client("<you>/sway")
html, raw = c.predict("My parcel never arrived", '{"late": {"type": "noul", "instructions": "Is the order late?"}}',
api_name="/decide")Pushing to main runs the tests and syncs the repo to the Space when the HF_TOKEN secret and the
HF_SPACE variable (for example yourname/sway) are set in the GitHub repo settings.
Laya by Convai Innovations, Apache 2.0. Game code MIT.