A guide to the community of Jev-like models. These are "System One" decision models: you send them a state and some typed questions, and they return calibrated choices, scores and yes/no probabilities instead of generated text.
Jev introduced the
/v1/systemoneinterface in September 2026. Within weeks it was served by a growing family of hosted and open models, including Mercury Decide, Liquid d1, Solar Decide, Laya, Kev, Decider and more.This list merges, de-duplicates, cross-checks and summarizes 106 community "awesome-jev" repositories.
18,330 unique links from 106 lists Β· 5,410 cited by 3 or more lists Β· 2,533 GitHub repos verified live Β· sources last pulled 2026-09-30 (sources.md)
β
= GitHub stars when verified on the pull date Β· πN = the number of the 106 source lists that cite the entry. π is the cross-list consensus signal: an entry that many independently curated lists include has been vetted many times.
- Jev-like models in 60 seconds
- Quick Start: Jev-like inference for free
- What 106 lists taught us: key insights
- Reference docs and SDKs
- Ecosystem
- Jev-like models: hosted, open and local Β· Framework and platform integrations Β· Community SDKs Β· MCP servers Β· Agent skills and plugins
- Coding-agent tooling Β· Browser, desktop and mobile Β· Search, RAG and data
- Evaluation, calibration and QA tooling Β· Games, robotics and simulation Β· Apps and domain applications
- Benchmarks and independent studies
- Research papers
- Articles, analysis and critique
- Learning resources
- Directories and other lists
- About this meta-list
Deep dives in this repo:
| docs/open-models.md | The Jev-like model family: open replications, local runtimes and hosted alternatives, with caveats |
| docs/patterns.md | Question design, composition patterns, thresholds, economics, security, a production checklist and anti-patterns |
| docs/evidence.md | Every independent benchmark, calibration audit, robustness probe and negative result, with numbers |
| docs/what-is-jev.md | A reference for Jev, the original model: specs, API, timeline, and contradictions between sources, resolved |
| docs/papers.md | ~35 papers on Jev and Jev-like models, plus the research lineage |
| docs/source-lists.md | A review of all 106 source lists: which to read for what, and which to treat with caution |
| catalog/ | The complete union of every link from every list, in 21 categories, ranked by consensus |
A Jev-like, or System One, model is not a chat model, and it never writes text. You send it two things in one request:
- a state: the evidence (a string, a JSON object, or an array of text);
- a map of typed questions about that state.
It answers every question in parallel, in a single call. Each answer is a probability distribution that your code can threshold.
Every System One model in the family speaks the same /v1/systemone wire format. Here is a support ticket with one question of each type:
POST /v1/systemone
{
"model": "<model id>",
"state": {
"ticket": "I was charged twice for my order last week and nobody has answered my emails. I want my money back today."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "Bugs, outages, errors",
"other": "Anything else"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Annoyed but civil", "Angry or threatening to leave"]
},
"wants_refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}The response contains one typed answer per question:
{
"model": "<resolved model version>",
"answers": {
"team": { "type": "choice", "choice": "billing",
"probabilities": { "billing": 0.95, "technical": 0.02, "other": 0.03 },
"confidence": 0.93 },
"frustration": { "type": "score", "score": 1.62,
"legend": { "0": "Calm", "1": "Annoyed but civil", "2": "Angry or threatening to leave" },
"probabilities": { "0": 0.03, "1": 0.32, "2": 0.65 }, "confidence": 0.71 },
"wants_refund": { "type": "noul", "noul": 0.97 }
},
"usage": { "input_tokens": 214, "output_tokens": 5 }
}Your code then decides what to do with the answers:
a = response["answers"]
if a["team"]["confidence"] >= 0.8: # "the answer is what; confidence is whether to act"
route_to(a["team"]["choice"])
else:
send_to_human_triage()
if a["wants_refund"]["noul"] > 0.9 and a["frustration"]["score"] >= 1.5:
escalate_priority()| Question type | Asks | You supply | You get back |
|---|---|---|---|
| Choice | "Which one?" | an instructions string and criteria as a map of option β description (up to 255 options; always include an other) |
choice, probabilities per option, confidence |
| Score | "Where on this ordered scale?" | an instructions string and criteria as an ordered list of 2β10 levels |
score (probability-weighted and 0-indexed, so it can fall between levels), legend, probabilities, confidence |
| Noul | "Is this true?" | an instructions string |
noul: the probability, from 0 to 1, that the answer is yes. There is no confidence field, and 0.5 means "can't tell", not "medium". |
β The reference for the original model and the question-design guide go deeper.
This is how to use a Jev-like (System One) model for free. Mercury Decide is Inception's System One decision model, served on OpenRouter as inception/mercury-decide:free. Details from its OpenRouter page:
- $0 input and $0 output. Free endpoints are rate-limited.
- 32,768-token context.
- About 0.42 s p50 latency and up to ~14 decisions per second.
- The same
/v1/systemonequestion schema shown above.
1. Get a free OpenRouter API key at openrouter.ai/settings/keys.
2. Call the Decisions API. Mercury Decide is a decisions model, so it uses OpenRouter's Decisions API, not /chat/completions. OpenAI-style chat SDKs won't work with it.
export OPENROUTER_API_KEY=sk-or-...
curl https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inception/mercury-decide:free",
"state": "I was charged twice for my order and want my money back.",
"questions": {
"team": { "type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": { "billing": "Charges and refunds", "technical": "Bugs and outages", "other": "Anything else" } },
"refund": { "type": "noul", "instructions": "Is the customer asking for a refund?" }
}
}'3. Or call it from Python (pip install requests; no other SDK is needed):
import os, requests
resp = requests.post(
"https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
json={
"model": "inception/mercury-decide:free",
"state": {"ticket": "I was charged twice for my order and want my money back."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Charges and refunds", "technical": "Bugs and outages",
"other": "Anything else"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["Can wait a week", "Handle within a day", "Handle within the hour"]},
},
},
timeout=30,
)
resp.raise_for_status()
answers = resp.json()["answers"]
print(answers["team"]["choice"], answers["team"]["confidence"], answers["refund"]["noul"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"]) # e.g. 1.05 {'0': 0.04, '1': 0.87, '2': 0.09}Tips:
- Already have code for another
/v1/systemoneclient? OpenRouter also accepts the same request body athttps://openrouter.ai/api/v1/systemone, so you can point an existing client's base URL athttps://openrouter.ai/api. - Batch questions. Put every independent question about a state into one request. It costs the same as a single question, so use the batching to stay inside the free rate limit.
- Graduate when you need more. Move to a paid decision model, or self-host an open one (Laya, Kev or Ollaya), when you need higher limits or local data. Most of these keep the same schema; see the model family.
- Evaluate on your own labelled examples before trusting thresholds. Calibration doesn't transfer between models, so a 0.8 cut-off tuned on one model means nothing on another.
The same lessons show up again and again across the lists. Most measurements were taken on Jev, the first and most-tested model, and are labelled as such. The lessons about how to build apply to the whole family. Numbers link to the underlying studies in docs/evidence.md.
- Jev-like models are a System One decision layer, not a model swap. The winning architecture is always the same: code builds the candidates, the decision model picks or scores, code acts. An LLM is called only for text or open reasoning. Examples:
- Browser agents let the model choose the element and use an LLM only to type.
- Games give the model the legal moves.
- Extraction lets regex find spans and the model select one.
- The 100Γ speedups are real only against chains of LLM calls. Jev's launch figures ("193.6Γ faster / 444.6Γ cheaper") are a best case against the slowest comparator; the average across eight setups is 97.8Γ / 149.2Γ. A single question is only ~1.5β3Γ faster than a small chat model. The order-of-magnitude wins come from batching many questions per call, where latency is flat up to ~25 questions and batching 13 questions was 12.2Γ cheaper, and from deleting LLM calls that never needed an LLM.
- Question design is the whole game.
- Splitting one question into atomic Nouls took a task from 48% β 98%, and phishing detection from 62.6% β 95.0%.
- Rewriting the criteria moved accuracy from 70% β 96%.
- Renaming options from
0/1tono/yescut AUC from .81 to .58. - Always include an "other/unknown" option. Removing it took one benchmark from 0.95 to 0.00.
- Keep computation in code. Counting (33β65%), arithmetic, dates, multi-hop reasoning and open-ended generation are documented weak spots. Jev's own jaggedness page lists them, and independent tests on open replicas agree. Use one Noul per item and sum in code.
- Confidence is a routing signal, not a guarantee.
- The top-confidence band is very reliable (48/48 correct at β₯ 0.9 in one audit).
- "Confidently wrong" still happens, and Choice and Score skew overconfident.
- Calibration does not transfer across datasets, languages, phrasings, model versions or models.
- Fit per-question thresholds on your own labels, and pin a model version.
- Cascades are the most consistent cost/accuracy win, with one catch. Accepting confident decisions and escalating the rest to an LLM matched frontier accuracy at ~ΒΌ the cost in several studies. But LLMs repeated ~96% of Jev's confident errors, so cascades mostly save money rather than add accuracy. Price the fallback: one production replay was 96% cheaper per call but 4% more expensive overall.
- It is not a security boundary. Typed output blocks format attacks, not decision manipulation. Blunt injected commands mostly fail. Evidence-shaped text flips decisions: fake approvals, editor's notes and fluent context flipped 61% of correct answers in JevOut, and 65β73% on open clones. Put deterministic rules first, keep humans on irreversible actions, and track taint.
- Audit what tools send. Hands-on audits found community tools leaking secrets, sending
.pemfiles and screenshots, exposing keys, and failing open on API errors. Check data egress and fail-closed behaviour before installing. - Coding agents are the dominant use case. The largest clusters across all lists are context compaction, model/effort routing, tool-call gates, "done" verification, and skill/rule pruning. Compaction is the most contested of these: several careful evaluations (Hermes Agent, jev-use) did not adopt it.
- Classical baselines are still strong. Trained small classifiers (bge-small + LR at 93% on Banking77, TF-IDF on spam) and regex often match or beat decision models. The durable edges of System One models are zero training, cost, latency, and robustness under distribution drift: on drifted spam, Jev held at 97.3% while TF-IDF fell to 72.5%.
- The family grew fast, but "compatible" β "equivalent". Within days, dozens of open models and
/v1/systemoneservers appeared (Laya, Kev, SemIf, NanoJev, Decider, Ollaya), followed by hosted alternatives (Mercury Decide, Liquid d1, Solar Decide). The shared schema makes them drop-ins for your code, not for each other's accuracy or calibration. Most "beats Jev" claims are in-distribution, so measure on your own data. - Most of the evidence is young. Nearly all numbers are author-reported, small-n, and from the family's first weeks, and the ~29 arXiv papers are unreplicated preprints. The best lists label every number (vendor / author-reported / independent), and this one does too.
Jev's documentation is the most complete public description of the /v1/systemone interface: its question types, confidence semantics, composition patterns and cookbooks. Because Jev-like models share the schema, most of it applies to every System One model.
Interface docs
- API reference
π31βPOST /v1/systemone, with request and answer shapes for all three primitives. The full docs index (llms.txt)π22is also available. - Concept pages:
- Primitives
π27 - Confidence
π30 - State
π10 - System One
π13 - How to build with System One
π17 - Use-case map
π17
- Primitives
- Model jaggedness: jev-1.13
π35β Read this before building. Nine documented failure modes (literal reading, math, dates, indirection, distracting state, adversarial content and more). They are written for Jev but are a good test plan for any Jev-like model. - Introducing System One Models and Jev
π57β The launch post that defined the category: RLCD (calibrated-decision training), the parallel sampler, and the Doom and Wikiracing demos. - Workflow evals
π32β Vendor evaluations on four workflows. The reference labels are averaged from two frontier models, not ground truth. The code is in WorkflowEvalsπ3.
Reference SDKs and tools
- Python: typesafe-sdk-python
β 257 Β· π59. JS/TS: typesafe-sdk-jsβ 260 Β· π57. Both are/v1/systemoneclients. Set the base-URL environment variable to point either one at any compatible server (local Kev, Decider or Ollaya; OpenRouter; and others). - system-one-adapter-python
β 368 Β· π61β The same client interface answered by an ordinary LLM. Use it for A/B baselines and fallback. - skills
β 2,491 Β· π68β An agent skill that teaches coding agents how to pick a question type and design questions. - For model-agnostic clients across many languages, see Community SDKs.
Patterns and cookbooks (the full annotated index covers all ~18)
- Patterns
π19:- Speculative fan-out
π16 - Confidence-gated routing
π15 - Composite scoring
π14 - Intent routing
π14
- Speculative fan-out
- Parallel questions
π23β 13 questions in one call: 12.2Γ cheaper and 10Γ faster. - Hierarchical classification
π15β Beam search past the 255-option cap. - Re-ranking
π15β Legal search: top-1 5% β 18%. - Guardrails for LLMs
π16 - Classifying RAG passages
π17 - Citation check
π18 - Pre-parsed value extraction
π14 - Date extraction
π16 - SDE cascade
π13 - Skill suggestion
π17 - Function calling
π17 - The OpenRouter cookbooks gate tool calls and verified cascade carry the same recipes over to models served by OpenRouter.
Hand-picked from the consensus of the source lists: mostly entries cited by many lists, plus a few high-signal ones that fewer lists noticed. Each section links to its full catalog page.
The System One model family itself. Many members serve /v1/systemone, so the same client code works across them. docs/open-models.md maps the landscape and its caveats.
Hosted
- Mercury Decide (Inception) β A structured decision model on OpenRouter, free, with a 32K context and the
/v1/systemoneschema. See the Quick Start. - Jev (TypeSafe) β The original and most-tested model, with a reference.
- Liquid AI d1 β Serves
/decisions/v1/systemoneand has a free tier. Reportedly #1 on the Jev Decision Index. - Upstage Solar Decide β Same schema, 512K context.
- Respan Span-01 β A System One model for behaviour monitoring.
- The OpenAI Decisions API (Luna, preview).
Open and local
- NandhaKishorM/laya
β 29,237 Β· π38β Convai's Apache-2.0 non-autoregressive encoder (421M/322M, 100+ languages, ~33 ms). Runtimes: laya-mlxβ 6,654 Β· π23(7β14 ms on M3 Max) and receptron/layaβ 661 Β· π20(Node/ONNX). - jaredpalmer/kev
β 8,057 Β· π57β A trainable family of System One models on Qwen3.5/3.8 with a pointer head and a/v1/systemoneserver. - TheoLeeCJ/SemIf-OpenJev
β 4,615 Β· π63β "Semantic ifs" from open models on a 3090, using shared-prefix logit readout. - TianyuCodings/NanoJev
β 2,453 Β· π60Β· vinnylarouge/jevlikeβ 1,335 Β· π57Β· Mapika/deciderβ 993 Β· π45Β· bespokelabsai/nimbleβ 1,981 Β· π21Β· wfzyx/vonβ 794 Β· π48β Trained replicas and recipes (0.4Bβ35B), several with published limits. - nokia-applied-research/AnyJev
β 986 Β· π32Β· featherless-ai/simple-jevβ 575 Β· π40Β· razorback16/openjevβ 548 Β· π49Β· ekzhang/openjev-sglangβ 335 Β· π48Β· githubnext/localjevβ 803 Β· π16β Turn any open LLM into a decision endpoint, with no training. AnyJev adds option-order correction. - ollaya-dev/ollaya
β 1,032 Β· π24β "Ollama for decision models": it runs open System One models locally. It serves Laya, Decider, NLI and GLiClass behind a/v1/systemone-compatible API. - Trackers: HF Jev Reproductions Tracker Β· Jev Decision Index
- β catalog/open-models.md
System One decision models landed inside mainstream frameworks within days of Jev's launch, usually as a new evaluate / decide / classify model type alongside generate. Most of these integrations target the shared /v1/systemone shape.
- vercel/ai
β 27,056 Β· π7β The AI SDK's decision-model provider maps Choice, Score and Boolean ontoexperimental_evaluate. - vercel/eve
β 5,426 Β· π28β Vercel's agent framework. A decision model (Jev by default) powers itsevaluatepath and its tool-approval policies. - vercel-labs/fx
β 3,234 Β· π9Β· vercel-labs/ai-cliβ 817 Β· π23β fx has a decision-model permission reviewer, reported at p95 18Γ faster than GPT Luna. ai-cli adds anevaluatecommand. - langchain-ai/langchain β A decision-model partner package: a classifier plus routing and auto-mode middleware. See also Building a harness with Jev.
- pydantic/pydantic-ai
β 20,296 Β· π11β A decision-model backend:output_typefields become typed questions (bool β Noul, Literal β Choice, IntEnum β Score). - BerriAI/litellm
β 59,945 Β· π9β A decision-model complexity router, a compaction guardrail, and a pass-through proxy. - ComposioHQ/composio
β 30,374 Β· π14β A decision-model provider for tool and argument selection. - BoundaryML/baml
β 9,363 Β· π5Β· BoundaryML/feelingsβ 22 Β· π10β Maps the type system onto the primitives.feelingsis a typed.feels()"AI if-statement". - Arize-ai/openinference
β 1,240 Β· π7β Records decision-model calls (Choice/Score/Noul) as OpenTelemetry spans. See also the Langfuse integration. - trycua/cua
β 27,561 Β· π17β A computer-use platform with a bounded-action recipe for decision models. - agentgateway/agentgateway
β 5,109 Β· π12β A decision-model guardrail for jailbreaks, harmful content and secret leakage at the gateway. - milvus-io/bootcamp
β 2,445 Β· π16β Nine search notebooks that use a decision model for ranking, filtering, routing and stopping. - zilliztech/GPTCache
β 8,206 Β· π7β Uses a Noul to check whether a cached answer serves a new request. - PrefectHQ/fastmcp β A two-stage MCP tool-search transform: a wide Choice, then a Noul per shortlisted tool.
- spring-ai-community/spring-ai-typesafe
β 41 Β· π20Β· crmne/ruby_llmβ 4,423 Β· π5Β· laravel/aiβ 1,203 Β· π6β Java/Spring, Ruby and Laravel support. - juspay/neurolink
β 143 Β· π23β Makesdecidea first-class inference type next to generate and stream, across 40 providers. - NousResearch/hermes-agent
β 250,326 Β· π5β A decision-model provider, skill routing, and a published (negative) compaction scorecard. - dubinc/dub
π3β A production link-safety gate ontypesafe-ai/jev. - More: TanStack AI (
decide()), DeepEval, MotherDuckprompt_jev(), OpenRouter cookbook: gate tool calls, Ollama 0.35 decision models Β· β catalog/sdk.md
Every major language had a community SDK within a week. Check the last commit date before you depend on one.
- Go:
- Stumble/jev-go
β 6 Β· π30β supports several gateways. - Gaurav-Gosain/jev-go
β 5 Β· π29 - mattn/go-jev
β 39 Β· π18β SDK and CLI. - Tangerg/typesafe-sdk-go
β 9 Β· π24
- Stumble/jev-go
- Rust:
- Twister915/typesafe-ai
β 13 Β· π33β async and blocking clients, observable retries. - gilljon/typesafe-ai-rs
β 6 Β· π31 - AbdelStark/s1-rs
β 1 Β· π23 - luizribeiro/jevrs
β 0 Β· π6
- Twister915/typesafe-ai
- Ruby:
- obie/ruby_decision_model
β 51 Β· π32 - joshmn/typesafe-sdk
β 7 Β· π35 - kieranklaassen/ruby_llm-typesafe
β 18 Β· π30 - carldaws/hunch
β 16 Β· π24β probabilistic control flow for Rails.
- obie/ruby_decision_model
- Elixir:
- dannote/jev
β 34 Β· π43β pattern-match decisions inside OTP. - nshkrdotcom/typesafe_sdk
β 5 Β· π36
- dannote/jev
- .NET: saibimajdi/typesafeai-dotnet-sdk
β 12 Β· π34Β· Hawxy/TypeSafeAI.Netβ 4 Β· π24 - Java: Premo-Cloud/typesafe-sdk-java
β 10 Β· π23 - Scala: jamesward/zio-typesafe-ai
β 5 Β· π29 - Swift:
- alterhq/typesafe-sdk-swift
β 4 Β· π25 - ainame/swift-typesafe
β 16 Β· π22 - peterfriese/system-one-foundation-models
β 54 Β· π20β maps Apple Foundation Models@Generabletypes onto the primitives.
- alterhq/typesafe-sdk-swift
- Haskell: inanna-malick/jev-dsl
β 7 Β· π20 - Python helpers:
- pithings/advocaat
β 96 Β· π48 - AboveColin/jevclient
β 2 Β· π30 - jlowin/vibecheck
β 39 Β· π8 - ktaletsk/jevframe
β 20 Β· π18β pandas and Polars. - docxology/daf-jev
β 6 Β· π27
- pithings/advocaat
- TypeScript helpers:
- jomatsu/zod-jev
β 8 Β· π23β semantic rules inside Zod schemas. - mateonunez/jod
β 3 Β· π22
- jomatsu/zod-jev
- CLIs:
- shiftynick/jev-axi
β 25 Β· π39 - tumf/jev-cli
β 13 Β· π32β a CLI and stdio MCP server. - sharziki/semdecide
β 75 Β· π37β Unix-pipeline and CI decisions with exit codes. - cristianoliveira/jeq
β 10 Β· π11β jq meets decision models.
- shiftynick/jev-axi
- β catalog/sdk.md
- jkudish/jev-mcp
β 469 Β· π66β The most-cited MCP server. Its typed tools (verify, screen for injection, find, rerank, classify, decide) follow a fail-closed contract. Install withclaude mcp add jev -- npx -y @jkudish/jev-mcp. - itsmostafa/system-one-connector
β 337 Β· π61β Formerlytypesafe-mcp. A single connector for Jev, d1, CLM and Laya. - blakestone-x/jev-mcp
β 25 Β· π42Β· burnigtm/jev-mcpβ 61 Β· π29Β· BYK/jev-mcpβ 3 Β· π20Β· PyModel/jev-judge-mcpβ 69 Β· π17Β· Brainwires/jevwireβ 21 Β· π28β Alternatives. BYK's server is eval-first, and jevwire adds an escalate-only Claude Code plugin. - kbhuw/jev-sift
β 47 Β· π18β Classify first, read selectively. Batch text classification for agents. - jiawei686/jev-ultrafast-mcp
β 18 Β· π17β Hands off a whole browser task in one MCP call. - agent-chaperone/agent-chaperone
β 21 Β· π21β An MCP proxy that screens both tool calls and tool results. - Note: several
jev-mcprepos share a name, so check the owner. β catalog/mcp.md
- dbreunig/building-with-jev-skill
β 145 Β· π39β The best community skill for writing decision-model programs: pick the primitive your code branches on, split two-property questions, and batch everything that shares a state. - kerpopule/hermes-jev-skills
β 932 Β· π42β The most complete decision-model-in-the-agent-loop pack: routing, memory, compaction, skill selection, and computer and browser use (Hermes, Claude Code, Codex). - wuyoscar/jev-skill
β 554 Β· source listβ Five skills, ajev-decideCLI, 108 agent-supervision scenarios, and honest evals, including a negative agent result. - Dicklesworthstone/skillranker
β 125 Β· π49Β· ShivamPansuriya/jev-skill-gateβ 7 Β· π24Β· GodsBoy/jev-agent-skill-routerβ 23 Β· π33Β· EliaAlberti/jev-rulesβ 62 Β· π27β Show the agent only the skills and rules the current turn needs, and allow "none". jev-skill-gate cuts the skill manifest by ~75%. - shitianfang/jev-use
β 33 Β· π34β Hands agent steps that need no text output to a System One model (p50 ~230 ms). Its compaction and agreement evals are published. - altryne/jevify
β 36 Β· π30Β· samtay32/jev-system-architectβ 2 Β· π23β Audit a codebase for fuzzy judgments that could become typed decision points. - aitofy-dev/jev-awesome-skills
β 2 Β· source listΒ· Pleo2/awesome-jev-agent-skillsβ 0 Β· source listβ Ready-made proceed/ask/stop gates, triage, diff review and QA-evidence skills. - β catalog/skills.md
The largest cluster in the ecosystem. It covers Claude Code, Codex, Pi, Hermes, OpenCode and Cursor.
Context compaction and pruning. This is contested, so read the evidence first.
- tamaratran/fast-jev-compaction
β 7,250 Β· π71β The most-starred project in the category. It replaces Claude Code's compaction summary with keep/drop decisions per tool call, and kept content stays verbatim. It began as a viral demo. - tamaratran/jev-pruner
β 153 Β· π35β Trims long Bash output before the model sees it. - GhalebDweikat/winnow
β 99 Β· π39β A calibrated context sieve with stub/recall pointers. It falls back to an LLM-backed adapter when the decision model is unreachable. - kunchenguid/compact-adviser
β 189 Β· π15β Asks "is this a safe moment to compact?" - joelhooks/pi-fast-jev-compaction
β 13 Β· π27Β· leonaaardob/fast-dev-compactionβ 9 Β· π26Β· compozy/yoshiβ 27 Β· π30β Ports to Pi and Codex, and a pruning proxy.
Model and effort routing. Route at session boundaries and run in shadow mode first.
- gargpratyush/jev-router
β 505 Β· π60β Routes Claude Code and Codex to the cheapest capable model. - 0xNatoshi/jev-codex-router
β 277 Β· π54β Per-turn model, reasoning-effort and speed routing for Codex (~60% savings on a 237-turn backtest). - BillionsBobby/JevRouter
β 309 Β· π42β One candidate set covering models, tools, subagents and skills. - ruban-24/switchboard
β 16 Β· π21Β· adarshmishra07/jcm-routerβ 8 Β· π27Β· xinyao27/jevonianβ 15 Β· π26β Cache-aware routing that leaves the cached main chat alone. - miuuyy/Astra-Ares
β 293 Β· π15Β· nidhi-singh02/agent-routerβ 99 Β· π25Β· prismhq/jev-routerβ 14 Β· π29β Mid-run effort selection, agent/CLI selection, and a router on LiteLLM.
Guardrails and tool-call gates. Put deterministic rules first, and choose fail-open or fail-closed explicitly.
- leepokai/jev-guard
β 48 Β· π38β Deny/ask/allow auto mode for any coding agent. It uses session context, flags injection in tool results, and tracks taint. Smoke test. - DevMortimer/pi-warden
β 153 Β· π50β Steers instead of interrupting. Rule breaks fell from 6 to 0 over 150 paired runs, and a replay on 17k calls held 42. - thruwire/foreman
β 618 Β· π61β A supervisor for Codex and OpenCode workers, with risk-tiered thresholds. - jomatsu/pi-jev-auto-mode
β 30 Β· π38β Auto-approves Bash, write and edit calls, and fails closed. - y0usaf/pi-jev
β 150 Β· π54Β· Nyarlathoteppppp/pi-heedβ 10 Β· π27Β· harshwasan/jev-sentinelβ 12 Β· π19β Pi extensions that judge whether an action is destructive, exfiltrates data, stays in scope, or matches what you asked for. - eugeniughelbur/jev-engineering
β 5 Β· π19Β· jesset/pi-verdictβ 10 Β· π14β Rules first, then a decision model for the grey zone. jev-engineering has a rerunnable 300-call injection test. - anpicasso/hermes-jev-approvals
β 19 Β· π27β Hermes command approvals: 8.7Γ faster, with 4.4Γ fewer prompts across 153 real commands. - luantak/is-malicious
β 32 Β· π27Β· hemanth/pkg-gateβ 0 Β· π16β Screen codebases and npm lifecycle scripts before you run them.
Done-verification, review and semantic lint
- devagrawal09/jev-review
β 646 Β· π74β Staged code review (risk β file profile β evidence β severity β reviewer routing) with a local dashboard. - valentynkit/jev-belay
β 19 Β· π41Β· qkal/Cannyβ 106 Β· π32Β· noplan-inc/limpetβ 5 Β· π24β Stop-hook gates that block an unverified "done". Canny's rule: "deterministic hooks decide, the model advises". - coldteadotai/abide
β 469 Β· π19β Enforces AGENTS.md rules with one Noul per rule. - valentynkit/jev-commit
β 12 Β· π39β A pre-commit check that the message matches the diff; it blocks only on leaked credentials. - NiazMorshed2007/jev-review
β 230 Β· π48Β· supercorp-ai/supercovβ 141 Β· π39Β· lakeday-org/perchβ 316 Β· π29Β· mizchi/jev-lintβ 113 Β· π23Β· doeixd/jev-prefβ 10 Β· π32β Continuous quality review, coverage, and linting for semantic rules that a parser can't check. - fatwang2/jev-review-action
β 2 Β· π19Β· yamadashy/jev-labeler-actionβ 0 Β· π5β GitHub Actions for submission review and issue/PR labels. - β catalog/dev.md Β· catalog/memory.md Β· catalog/routing.md Β· catalog/security.md
The shared recipe: build a numbered list of what's on the page or screen (accessibility tree, DOM or OCR), let the decision model pick the operation and target, let code execute, and call an LLM only when text must be typed. Most decision models are text-only and can't see pixels.
- browser-use/jev-ultrafast
β 21,563 Β· π76β The most-cited project in the entire ecosystem. It picks the operation and element in one request and booked a ZΓΌrichβLondon flight search in 7.1 s for $0.0039. The launch demo has been retested: 1/20 on complex tasks. - awlevin/typesafe-computer-use
β 1,098 Β· π46β macOS computer use: OCR the screen, then a decision model picks the action, for ~$0.0002 per step. - droidrun/mobile-jev
β 426 Β· π59β Android automation; an Uber booking took 9 actions and ~21 s. - moritzkremb/jev-voice-browser
β 374 Β· π49β Voice control at ~300 ms per spoken word. URLs are copied from your words, never generated. - jkudish/jev-browser
β 297 Β· π49Β· Ying-Kai-Liao/jev-browserβ 93 Β· π43Β· wy-coliney/jev-browser-useβ 721 Β· π45β In these, the LLM plans and the decision model decides. Ying-Kai-Liao's version solved 40/42 live tasks using ~70Γ fewer tokens. - Sac-Y/Jev-cu
β 613 Β· π28Β· lahfir/agent-desktopβ 1,733 Β· π16Β· savka777/jev-useβ 111 Β· π22β Desktop control through accessibility trees, with local policy gates on sensitive clicks. - kitze/unclutter
β 343 Β· π49Β· realZachi/typesafe-adblockβ 87 Β· π46Β· anishfn/shapeshiftβ 755 Β· π23β Browser extensions and UI: clutter removal, "is this DOM element an ad?", and a text box that morphs into the UI you mean. - β catalog/browser.md
- superagents-lab/jev-search
β 495 Β· π64β Web search where a decision model does source selection, query understanding and relevance ranking. It returns links, not generated answers. - realZachi/pg-jev
β 386 Β· π59β Ask Postgres tables questions in plain language. Its batching study found 20 rows per request gave 100%, and 80 rows gave 77β94%. - kylemclaren/jevql
β 14 Β· π43Β· colliber/duckdb-jevβ 27 Β· π19Β· mattn/sqlite3-jevβ 3 Β· π15Β· giuliosmall/pg_typesafeβ 87 Β· π34β Semantic SQL predicates. Every row leaves your machine, so pre-filter with cheap predicates first. - jexp/neo4jev
β 153 Β· π52β Graph navigation: a Choice picks the next relationship and a Noul asks "goal reached?", with beam search in the app. - dzhng/jevgrep
β 1,902 Β· π23Β· keltokhy/jgrepβ 131 Β· π32Β· uehaj/sys1grepβ 144 Β· π36Β· sufianetaouil/everyβ 7 Β· π32Β· ellipsis-dev/blinkβ 92 Β· π45β "grep by meaning". These send code to the API, so check what leaves your machine. - jerryjliu/docjev
β 487 Β· π31β Fast document classification and packet splitting, from LlamaIndex's founder. - AkashPriyadarshii/jev-curate
β 92 Β· π45Β· RenaGao/jev-dataopsβ 60 Β· π18β Dataset sifting (Rust, Parquet/JSONL) and streaming data selection with LoRA training. - WiktorB2004/llama-index-jev
β 8 Β· π33Β· hev/rerankerβ 14 Β· π25Β· hotchpotch/jev-rerankerβ 36 Β· π21β Rerankers. Fuse with BM25 or embeddings rather than reranking with the decision model alone. - keltokhy/jsort
β 24 Β· π24Β· keltokhy/jlinkβ 6 Β· π21Β· keltokhy/jselectβ 3 Β· π16β Sort by meaning (pairwise comparisons plus Bradley-Terry), record linkage, and budgeted evidence selection. - reachjalil/jevlogs
β 16 Β· π33Β· hyperspaceai/jevcacheβ 76 Β· π17β OpenTelemetry log triage, and a decision cache for deterministic CI replay. - β catalog/data.md
- abhixhek/jevcal
β 10 Β· π41β Stop guessing thresholds. It fits per-question thresholds on your labels, verifies them on held-out data, and fails CI when a model update breaks them. Cited in more "best practice" sections than any other tool. - sutro-sh/jev-align
β 300 Β· π42β Builds calibrated "AI functions" from human feedback, using active labelling plus GEPA to optimize the questions. - AbdelStark/jev-benchmarks
β 21 Β· π33Β· jmanhype/jev-dspy-labβ 10 Β· π22β Calibration, selective risk, and record/replay in DSPy. - openlayer-ai/jevals
β 98 Β· π26β Agent evals and guardrails as typed decisions (it also runs locally with Kev or Laya): eight judgments per trace for ~$0.00006. - smkrv/jev-calibrate
β 31 Β· π17Β· nikkoxgonzales/jev-certifyβ 1 Β· π11Β· sathariels/jevcheckβ 2 Β· π7Β· vcjdeboer/jev-reliabilityπ4β Criteria tuning, conformal routing bounds, behavioural contract tests that refusejev-latest, and repeatability/phrasing preflights. - suraj-phanindra/wellposed
β 2 Β· π17Β· ariel-frischer/jevkitβ 3 Β· π22Β· simota/tenbinβ 4 Β· π18β Lint your questions offline, before you pay for calls. - AntonioCoppe/jev-harness
β 15 Β· π31β Confidence gates, shadow mode, recipes and evals (on a row-filter job, Claude CLI took 48.9 s and the decision model 1.3 s). - mandu5/jevcompat
β 0 Β· π6β A conformance suite for/v1/systemone-compatible servers. Use it to check whether a Jev-like model really is a drop-in. - β catalog/bench.md
The recipe for games and control: feed structured state (RAM β JSON, legal-move lists), never pixels. Let code own the physics, the route and the arithmetic, and let the decision model pick at branches. Keep a fast deterministic reflex layer that can veto it.
- fhshaik/typesafe-mario
β 422 Β· π64β Super Mario Bros. from structured emulator RAM. It is the canonical "state as data" example. - RomanSlack/jev-drone
β 228 Β· π66β A MuJoCo drone. The decision model advises at 2.5 Hz while a 50 Hz reflex layer holds a veto. - valentynkit/jev-plays-pokemon-red
β 9 Β· π38β Code owns the route and the arithmetic, the model picks at branches in ~100 ms, and calibration is measured with Brier scores. - dabit3/jev-experiments
β 397 Β· π38β 22 latency-focused demo apps by Nader Dabit. - phyous/tsai-sc
β 27 Β· π45Β· sorrycc/typesafe-snakeβ 23 Β· π32Β· standardagents/jevpilotβ 204 Β· π32Β· rmalde/minecraft-agentβ 568 Β· π18β StarCraft, Snake (legal moves generated in code), a driving autopilot, and Minecraft (an LLM plans, the decision model acts). - FBddcz/embodied-jev
β 249 Β· π30Β· Dimweaker/jev-liberoβ 78 Β· π23Β· TarunTomar122/jev-askable-armβ 11 Β· π27Β· openroboto-ai/jev-robot-controlβ 53 Β· π13β Robot arms that pick primitives from menus, never torques. OpenRoboto compares Jev with GPT-6 Astra on cost and time. - β catalog/games.md Β· Frank-ZY-Dou/awesome-jev covers robotics in depth.
- AboveColin/HA-Jev
β 68 Β· π56β Home Assistant: ask your house a question and get a number back, as sensors and automation actions. - jarrodwatts/jev-trader
β 2,708 Β· π71β One trade decision per Monad block (~81 ms). It's educational: no list reports sustained trading profit, and bots should be dry-run by default. - kyotofin/tax-doc-classifier
β 483 Β· π41β IRS form pages: 100% strict on 261 forms at ~$0.001 per page, replacing production Sonnet. - fazlerocks/jevmail
β 92 Β· π27Β· parth-kp/jev-mail-classifierβ 19 Β· π19β Gmail triage: 1,000 emails in ~1 minute for ~3 cents. - jev-chat/jev-chat-jarvis
β 7,181 Β· π37β A read-only chat copilot for phones. An LLM drafts replies, a decision model ranks them, and you choose whether to send. - socai-io/jev-social
β 130 Β· π42Β· AkashPriyadarshii/jev-seoβ 88 Β· π32Β· usenotra/notraβ 225 Β· π19β Social research, SEO/GEO audits, and a production GEO platform. - brainstormity/Jev-Moderation-Bot
β 47 Β· π36Β· backmeupplz/jev_antispam_botβ 13 Β· π20β Discord and Telegram moderation. - OpenByteInc/QuantDinger
β 12,344 Β· π29β An AI trading OS that uses a decision model as an evidence and risk gate. It blocks entries but lets exits and protective orders through. - choxos/jev-reviewer
β 37 Β· π22Β· youkiti/tiab-review-plugin β Systematic-review screening (95% recall on 16,645 records) and extraction. - monteduro/killmyidea
β 244 Β· π33Β· ChetasLua/jevmeterβ 102 Β· π45β Fun ones: kill, fix or ship your startup idea, and a live decision meter on any video. - β catalog/apps.md Β· catalog/finance.md Β· catalog/agents.md
Summarized with numbers in docs/evidence.md.
- punk2898/awesome-jev-verified β The best claim-by-claim audit. An open 2,390-question benchmark with per-question logs, run against four GPT models. It confirms the price, but not "sub-100 ms", "445Γ cheaper" or "eliminates hallucination".
- fstandhartinger/jevbench
β 189 Β· π40β JevBench: 534 frozen, half-sealed decisions and a four-axis composite across 40+ systems. See also the Benchmark Heaven leaderboard. - Jevals.com
π12β Jev vs six LLMs on human-labelled sets, with the data published. - OpenRouter: Is Jev as accurate as frontier models? β Banking77: 81.0% vs Opus 5's 84.4%, at 175 ms vs 2,266 ms and $0.11 vs $2.42 per 1k.
- willkelly/jev-evaluation
β 0 Β· π11β 123,805 pre-registered requests. ECE was 0.075 in-domain, but calibration fails OOD and authority injection works. - anisselbd/jev-phishing-bench
β 7 Β· π32β The canonical question-atomization lesson: 62.6% with one question vs 95.0% atomized. Haiku 4.5 wins the direct question. - ickma2311/jev-baselines-eval
β 5 Β· π12Β· bitnovus/jev-spam-evalβ 2 Β· π34β Classical baselines win in-domain, but Jev wins under drift. - anessbelbati/jev-rerank-bench
β 9 Β· π39Β· zhuyansen/jev-search-rerank-evalβ 9 Β· π31β Reranking: Jev ties Cohere, but doesn't beat embeddings alone, so fuse them. - jujumilk3/jev-calibration-audit
β 0 Β· π15Β· KantaHayashiAI/jev-does-not-play-diceβ 3 Β· π17Β· scienthoon/jev-ood-calibrationβ 6 Β· π25Β· jourdanlabs/assay-001β 0 Β· π18β Calibration audits covering abstention, known probabilities, OOD, and a pre-registered test. - Gaurav-Gosain/jev-sec-bench
β 3 Β· π31Β· zkousama/jaggedπ4β Prompt injection and vulnerable-code detection. - nekuda-ai/WindTunnel
β 86 Β· π18β WebMCP: Jev + Mercury solved 49/49 tasks, bare Jev 25/49. - Zaious/jev-capability-atlas
β 26 Β· π27Β· RINNECODER/jev-behavior-studyβ 3 Β· π23Β· FirasSX914/Janusβ 2 Β· π25β Where Jev holds up vs breaks, order and lost-in-the-middle effects, and when routing is worth it. - thevibeworks lab Β· rhc98 calibration report Β· wuyoscar evals β First-hand measurements published inside source lists.
- Write-ups:
- agentjournal.dev
π21: one judge call vs dimension scores. - amankumar.ai: 16,000 calls; "a filter, not a replacement".
- primeline.cc: pre-registered; 4 of 12 failure modes real.
- Every's vibe check
π19 - Lindfors
- Near Here
π14 - YTAL production replay: not adopted.
- agentjournal.dev
- β catalog/bench.md
~35 papers on Jev and other System One models appeared within two weeks, all unreplicated preprints. The full annotated list is in docs/papers.md.
- Jev in the Wild
π18β An ecosystem survey of 2,170 GitHub projects. - JEV-as-a-Judge
π7Β· JEV vs. LLMs as Rubric Judgesπ5β Jev comes within ~3 pts of GPT-6 at ~0.36% of the cost, but LLMs repeat its confident errors, which caps cascade gains. - Type-Safe Is Not Error-Free
π5Β· JevOutπ4Β· JevAdvBenchπ5Β· Decision Hijackingπ6β Robustness: option names, natural context and injection. - REFLEX
π4Β· Jev-Mobileπ4Β· JEV-Star Β· Jev-Memπ6β Planner/executor agents that make 66β73% fewer strong-model calls. - Evaluating Decision Models for Text Annotation in CSS Β· Calibrated Decisions at Scale: crash narratives
π4Β· Just Ask Jevπ7β Annotation, large-scale coding and alignment-failure detection. - Typed Decision Models: An Early Evidence Audit β Across 28 papers there is no independent accuracy advantage; the gains are latency and cost.
Analysis
- Simon Willison: Jev introduces a new shape of LLM β An influential early analysis.
- Latent Space: Jev, System One models for Prod, Not God β A founder interview (2h21m). Also: a System One model that only decides
π11and 6 clones of Jev in 2 days. - Sean Goedecke: Jev means structured output is interesting again β A skeptical technical take. A one-constrained-token baseline gets 2β3Γ speedups.
- Warmer Sun: Typed decisions, not chat β Separates the launch claims from the vendor's own footnotes.
- Flavio Copes: A deep dive into Jev
π17β Often named as the best long-form introduction. - Pere Pages: Jev sorted Β· sgnt.ai: You could have built Jev Β· Archer Hume: Jev's architecture unmasked β Claim-by-claim decoding, the single-token-logit recipe, and an architecture inferred from ~10k API calls.
- Anthony Maio: The language model that won't talk
π8Β· TrueFoundry: What actually shipped Β· Silverthread Labs: What the 193Γ benchmark doesn't measure Β· ContextOS: production engineering review β Critiques. - Sebastian Raschka: From bag-of-words to Jev Β· Vercel: What is Jev? Β· OpenRouter: Jev vs LLM, when to use each β Context and positioning.
News
- TechCrunch
- The Register: plays Doom Β· The Register: Shut up and calculate
- Business Wire: $40M seed
- Forkast: Not an LLM, and that may be the point
- The New Stack: OpenAI's Decision API
- Chinese coverage: 36Kr Β· Huxiu hands-on
Community discussion
- HN launch thread
- Ask HN: Noul as a decision primitive
- "It's the inference technique, not the training"
- "The threshold is part of the prompt"
- tanxarx/awesome-jev collects ~65 skeptical threads with numbers.
- β catalog/community.md (3,000+ threads and posts)
- Li-Evan/awesome-jev: cheatsheet β The best one-page technical reference, covering endpoint, schemas, limits, jaggedness workarounds and composition moves. Li-Evan also publishes a free Chinese book, Jev in Practice.
- disler/ten-levels-of-jev
β 124 Β· π10β From one smart if-statement to a coding agent that reaches for a decision model on its own. - harshithsunku/learn-jev-end-to-end
β 14 Β· π15β A free hands-on course: 13 agent use cases, a fast brain and a slow brain, one OpenRouter key. - nexibeo/jev-cookbook
β 34 Β· π27Β· datawhalechina/jev-cookbookβ 46 Β· π14Β· paramjeetn/jev-cookbookβ 7 Β· π11β Tested recipes on OpenRouter, 18 Chinese notebook recipes with evals and local fine-tuning, and 120+ use cases. - dog-last/awesome-jev β A "should you use a decision model?" decision tree and cookbooks tested in CI against the live API.
- Justmalhar/awesome-jev-apps β 100 runnable apps, plus an app spec that lists what never to ask the model.
- vicfei/awesome-jev-prompts β 43 question-design patterns and 10 anti-patterns.
- davila7/jev-explained
β 33 Β· π14Β· Foadsf/jev-for-engineersβ 5 Β· π23Β· PromptEngineer48/langchain-jev-tutorial β An illustrated primitives guide, engineering examples, and a LangChain support-ops agent. - Articles and videos:
- Chinese and Japanese:
- Bald0Wang/jev-docs-zh: Chinese translation of the docs.
- yibie/jev-engineering-zh: the founder's design notes in Chinese.
- kuhung/understanding-jev
- yzfly/awesome-jev-zh
- mukishitsuu-png/awesome-jev-ja
- β catalog/learn.md
- Live directories:
- awesomejev.com
π9 - madewithjev.com
π14: builds with reported cost and speed. - jevusers.com
- jevlist.ai
- mrjev.com: hands-on security reviews.
- jevforagents.com
- awesomejev.cc
π4 - logicrw radar
- systemonemodels.org
- awesomejev.com
- The 106 source lists: sources.md has stars, pull dates and commits. docs/source-lists.md says which list to read for what, and which to treat with caution.
- Most-starred sources:
- Specialized sources:
- Robustness: Yifan-Lan/awesome-jev-robustness
- Security papers: Sarim-MBZUAI/awesome-jev-security
- Papers: OmniJev/awesome-jev-papers
- Open models: andyrewlee/awesome-system-one
- Robotics: Frank-ZY-Dou/awesome-jev
How it was built
- Pull. All 106 source repositories listed in sources.md were shallow-cloned on 2026-09-30. Their star counts and commits were recorded.
- Union. Every Markdown link in every file was extracted: 127k link occurrences, which de-duplicate to 18,330 unique entries after URL normalization. GitHub deep links collapse to
owner/repo, and links to the source lists themselves are dropped. - Consensus. For each entry,
πcounts how many distinct lists cite it. - Verification. The 2,533 GitHub entries cited by β₯5 lists were fetched live to confirm they exist, follow renames, and record stars and descriptions. The 21 that didn't resolve are listed in catalog/unresolved.md.
- Categorization. Entries were categorized automatically for the catalog. The sections above were hand-curated.
- Synthesis. Every source list was read in full. Per-source notes are in docs/source-notes/. The insights, the reconciled facts and the evidence digest in docs/ were synthesized from those notes.
Caveats
- This is a snapshot of a two-week-old ecosystem. Specs, prices and access change quickly, so check the official models page.
- Most numbers are author-reported and unreplicated. They are labelled as such where it matters.
- Inclusion is not endorsement. Many launch-week repos are thin scaffolds, and many tools send your data to third parties.
- This list is not affiliated with any model vendor.
Updating
scripts/refresh.sh # re-pull all sources β stars/commits β catalog β sources.md β README numbersTo add a source, append its GitHub URL to sources.md and refresh. The hand-written synthesis in docs/ and templates/README.md.tmpl should then be reviewed against git diff catalog/.
Repository layout
README.md β generated from templates/README.md.tmpl (live β
/π numbers)
sources.md β the 106 sources: stars, pull date, commit
docs/ β synthesis: what-is-jev, patterns, evidence, open-models, papers, source-lists
docs/source-notes/ β one review note per source list
catalog/ β the complete categorized union (18,330 entries)
data/ β catalog.json/csv, sources.json, verified.tsv, source_profiles.json
scripts/ β pull_sources, extract_links, build_catalog, verify_github, render_*, refresh.sh