Skip to content
CallSohailPublic

About

A word game against the Laya decision engine

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

10 Commits

Folders and files

Repository files navigation

title SWAY
emoji 🎯
colorFrom blue
colorTo gray
sdk gradio
sdk_version 6.28.0
python_version 3.12
app_file app.py
pinned false
license mit
short_description A word game against the Laya decision engine
models
convaiinnovations/laya

SWAY

A word game played against Laya, an open non-autoregressive decision engine. Each case sets a goal on the engine's answers. You write the message, the engine judges it in one forward pass, and you see every probability it produced.

The cases are real production workflows: support triage, content moderation, LLM guardrails, model routing and email security. Playing them is a hands-on way to see what typed decisions, calibrated confidence and language routing do.

The cases

# Case Goal
1 First contact Get routed to billing without billing words
2 Cold anger Sound clearly annoyed while staying non-toxic
3 The quiet exit Trigger churn risk while sounding calm
4 Tick tock Convey urgency with no alarm words
5 Babel desk Report an outage in a language other than English
6 Fog machine Make billing and technical tie
7 Gatekeeper Ask a real security question the guardrail does not flag
8 Short and hard Get rated hard in 70 characters or fewer
9 Real, not phish Write an urgent IT email that is not flagged as phishing
10 Grand finale A calm non-English refund request with no churn and no toxicity

Five shots per case. A shot blocked by a rule (banned word, length, language) is free. Scores reward passing with margin; three stars need a strong pass on the first shot. Every case carries a hint; using it caps that case at two stars. Progress is kept in the browser.

Each case also names the product it comes from, so the case card says where that exact decision runs.

The Lab tab runs any preset or custom question set on any input and shows the full answer.

Autoplay

The Autoplay tab plays the game by itself, with Laya making every decision. Pick one case or all of them, press start, and watch the search stream.

For each case the player composes eight candidate messages from a library of phrase fragments, runs them through the same sanitiser and rule checks a human shot goes through, and drops the blocked ones for free. The survivors are scored by the engine. It keeps the best three and mutates them: swap a fragment, drop a clause, add a closing line. It stops at the first candidate scoring 95 or more, or when the case runs out of model calls.

Setting Meaning
Quick 12 model calls per case
Thorough 30 model calls per case
Seed Same seed, same search. Default 42

A whole run stops at 300 model calls or six minutes, whichever comes first. The final card reports cases solved, average best score, model calls, average decision time and total time, and compares the machine's best scores to yours. Your own progress is only read, never changed.

Non-English cases get French and Spanish fragments, so a candidate is never a mix of two languages and the router always has a full sentence to detect.

Where these decisions run

Use case What the engine decides
Support triage Which queue a ticket belongs to, and how frustrated the customer sounds
Moderation Whether a post is toxic, in the same pass that reads it for anything else
LLM guardrails Whether a prompt is a jailbreak or an injection, before the model call
Model routing How hard a request is, so easy ones go to a small model
Email security Whether an urgent mail is genuine or phishing
Multilingual routing Which checkpoint answers, from script and language detected first
Confidence-gated automation Act above a threshold, send everything under it to a person

Run locally

python -m venv .venv && source .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cpu   # or your CUDA build
pip install -r requirements-dev.txt
python app.py            # http://localhost:7860

The first start downloads the checkpoints (about 1.5 GB for English plus multilingual).

Configuration

Variable Default Meaning
LAYA_CHECKPOINTS english,multilingual Checkpoints to load. multilingual alone runs in about 2.5 GB RAM
LAYA_DEVICE auto cpu, cuda or mps
LAYA_LOW_MEMORY 1 Build models on the meta device to cut peak RAM during load
SWAY_CONCURRENCY 2 Parallel model calls
SWAY_SKIP_ENGINE unset 1 starts the UI without loading models (CI)

Tests

pytest -q                                                       # fast, no model
SWAY_MODEL_TESTS=1 pytest -q tests/test_levels_model.py         # every case has a known passing answer,
                                                                # and Quick autoplay solves at least 6 of 10

API

The Lab is exposed as /decide:

from gradio_client import Client
c = Client("<you>/sway")
html, raw = c.predict("My parcel never arrived", '{"late": {"type": "noul", "instructions": "Is the order late?"}}',
                      api_name="/decide")

Deploy

Pushing to main runs the tests and syncs the repo to the Space when the HF_TOKEN secret and the HF_SPACE variable (for example yourname/sway) are set in the GitHub repo settings.

Credits

Laya by Convai Innovations, Apache 2.0. Game code MIT.

About

A word game against the Laya decision engine

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages