Calibrated yes/no, pick-one and scale judgments from your own LLMs — on text and on images — read straight off the logprobs.
A lot of code needs a small judgment that plain logic can't express: is this message asking for a refund?
Which team should handle this ticket? How severe is this alert? Does this new fact replace the old one?
The usual answer is a pile of if/else branches and regexes, or a chat prompt whose free-text answer you then
parse. Hunch replaces both with one typed call. You describe the check, and it returns a probability your code
can threshold, not a paragraph.
It works on photos too. Does this photo show flooding? Does it show the event in the headline? Is this a protest or a concert crowd? Send the image with the check and get a probability you can gate at 0.9, the same as for text. On 188 hand-labelled look-alike photos (a flooded street vs a wet one, earthquake rubble vs a demolition site, a building on fire vs one lit by fireworks), four vision models qualify, one of them a 4-bit 35B MoE on an edge-class GPU, with Gemma-4-31B at 98.4% and calibration error 0.016. See Images.
- Typed checks:
yesno→p_yes,pick(one of up to 300 options) → the best option and a full distribution,scale(up to 10 ordered levels) → an expected value and a distribution. Every answer carries aconfidence. - No text generation. Each check is a single constrained output token (
max_tokens: 1), and the answer is read from the token logprobs and renormalised over the allowed labels. The model can't ramble, and a malformed answer can't make it into your code. - Images too. Send photos with the request and ask about them (does this photo show flooding? does it show the event in the headline?) on any vision model your vLLM serves. Same checks, same probabilities, same gate: four models qualify on a hand-labelled set of 188 look-alike photos (see Images).
- Parallel by default. All checks in a request run concurrently, with the context first in every prompt so the server's prefix cache is reused.
- Your models, your hardware. Hunch runs against any vLLM OpenAI-compatible server. No data leaves your network.
Inside an agent that already talks to a vLLM endpoint, there's nothing to set up. Hunch reads the endpoint and
model the agent already uses (ANTHROPIC_BASE_URL / ANTHROPIC_MODEL, OPENAI_BASE_URL / OPENAI_MODEL, …) and runs
in-process, with no server and no config file:
from hunch import judge # pip install "hunch @ git+https://github.com/ihubanov/hunch"
r = judge({"ticket": "Charged twice for order A-104."},
{"refund": {"kind": "yesno", "question": "Is the customer asking for money back?"}})
r["results"]["refund"]["p_yes"]The same from a shell, or as a tool the agent starts itself (see Zero-config):
python -m hunch judge '{"context": "...", "checks": {...}}' # one call, JSON out
claude mcp add hunch -- python -m hunch mcp # a `judge` tool for Claude CodeAs a shared service:
pip install . # or: docker build -t hunch .
export HUNCH_BACKEND_URL=http://localhost:8000 # your vLLM server
export HUNCH_BACKEND_MODEL=Qwen/Qwen3.5-397B-A17B # a model it serves
python -m hunch selftest # must print SELFTEST PASSED
python -m hunch # serves http://127.0.0.1:8791curl -s localhost:8791/v1/judge -H 'content-type: application/json' -d '{
"context": {"ticket": "Charged twice for order A-104. Fix it today or I cancel."},
"checks": {
"refund": {"kind": "yesno", "question": "Is the customer asking for money back?",
"yes_if": "asks for a refund or reversal of a charge",
"no_if": "anything else, including only reporting a problem"},
"team": {"kind": "pick", "question": "Which team handles this?",
"options": {"billing": "charges, refunds", "technical": "bugs", "sales": "pricing"}},
"anger": {"kind": "scale", "question": "How angry is the customer?",
"levels": ["calm", "annoyed", "furious"]}
}
}'{
"model": "qwen3.5",
"results": {
"refund": {"kind": "yesno", "p_yes": 0.5927},
"team": {"kind": "pick", "pick": "billing", "probs": {"billing": 0.9997, "technical": 0.0002, "sales": 0.0001}, "confidence": 0.9975},
"anger": {"kind": "scale", "value": 1.647, "probs": [0.0027, 0.3477, 0.6496], "confidence": 0.3963}
},
"usage": {"prompt_tokens": 465, "completion_tokens": 3, "backend_calls": 3}
}That's a real response from Qwen3.5-397B (lightly trimmed). p_yes is 0.59 because the customer reports a double charge but never
explicitly asks for money back. Hunch reports that uncertainty instead of guessing, and your code decides what 0.59 means.
Not every model makes a good Hunch backend. In our tests the older Qwen3 generation was confidently wrong, and spelling out definitions made it worse. Check before you rely on a model:
python -m hunch qualify qwen # one or more model names from hunch.toml (default: all)It runs straight against your backend (no Hunch server needed) and uses the 240 fictional look-alike pairs that ship
with Hunch (hunch/lookalikes.py): two runs with the look-alikes named, plus one with the
bare question. That's about 720 one-token calls, roughly a minute at the default concurrency. A model is
QUALIFIED when:
| Criterion | Default | Why |
|---|---|---|
| Backend setup | constraint enforced, logprobs returned | otherwise answers are unconstrained |
Accuracy at p_yes ≥ 0.9 |
≥ 90% (--min-accuracy) |
the gate you'd actually use |
| Calibration error (ECE) | ≤ 0.15 (--max-ece) |
a wrong answer mustn't look as certain as a right one |
| Definitions don't hurt | named ≥ question-only − 1 point (--max-definitions-drop) |
if better-written checks make it worse, you can't fix it by writing better checks. A drop of a point or less is run-to-run noise |
| Stability | ≤ 2% of answers flip between identical runs (--max-flip-rate) |
the same input should get the same decision |
A backend that is down is reported as UNAVAILABLE, not as a failed model. The exit code is 0 only if every
model qualifies, and --json FILE writes the full report. --quick does one named run (no stability check).
python -m hunch selftest is the 10-second version (setup plus 12 cases). Run it after every deployment. Use
qualify when choosing or upgrading a model.
| Field | Type | |
|---|---|---|
context |
string, object or array | The data to judge. Use an object with named fields when there are several parts, and refer to them in questions with backticks, e.g. `ticket.messages[0]` |
checks |
object: id → check | Your own ids; results come back under the same ids |
images |
list of strings, optional | Up to 8 images for a vision model: https://… URLs the backend can fetch, or data:image/…;base64,… URLs. They are part of the context, numbered IMAGE 1…IMAGE n so questions can refer to them. context may be omitted when only images are judged |
model |
string, optional | One of the configured models; defaults to service.default_model |
effort |
string, optional | Thinking effort for a deliberate model on this call only (low / high / max). Ignored for one_token models (it would switch thinking on) and by backends without it. In hunch.toml, effort requires mode = "deliberate" or "auto" |
Check kinds:
kind |
Fields | Result |
|---|---|---|
yesno |
question, optional yes_if, no_if |
p_yes (0–1) |
pick |
question, options: key → description or null (1–300 options) |
pick (most likely key), probs (key → p), confidence |
scale |
question, levels: list of descriptions, lowest first (1–10) |
value (expected level, can fall between levels), probs (list), confidence |
confidence = 1 − normalised entropy of the distribution: 1.0 means all probability is on one answer, 0.0 means uniform.
Errors come back as {"error": {"code": "...", "message": "..."}}:
| HTTP | code |
When |
|---|---|---|
| 400 | invalid_request, unknown_model, too_many_options, too_many_levels, too_many_images, images_not_supported (images sent to a text-only model) |
Bad input (unknown fields are rejected) |
| 401 | unauthorized |
HUNCH_API_KEYS is set and the bearer key is missing or wrong |
| 413 | context_too_long |
The context plus the question exceed the model's context window |
| 502 | backend_error |
The backend returned an error, or did not enforce the constraint (never turned into a made-up probability) |
| 503 | backend_unavailable |
Still failing after retries (429 / 5xx / timeouts, with Retry-After honoured) |
GET /v1/models lists the configured models, each with the mode its checks actually run in (one_token or
deliberate; auto is resolved, with the configured value in configured_mode). GET /health checks that the backend serves them.
The common way to get a decision out of an LLM is to ask for JSON ({"is_refund": true}) and parse it. That works,
but you only get the model's top answer, with no sense of how close the call was. You pay for the output tokens,
reasoning models may think for hundreds of tokens first, and every so often the output isn't valid.
Hunch asks for exactly one constrained token and reads the probability the model put on each allowed answer:
| Ask for JSON | Hunch | |
|---|---|---|
| What you get | one answer (true) |
a probability (p_yes: 0.93) plus a distribution for picks and scales |
| Close calls | invisible: 51% and 99% look the same | visible, so you can route 0.4–0.6 to a human or a bigger model |
| Output tokens | tens to hundreds (more with thinking) | 1 |
| Malformed output | possible, needs retries or repair | impossible: the grammar only allows the labels |
| Thresholds | fixed by the model's wording | yours: act at 0.9, review at 0.5–0.9, drop below 0.5 |
| Many questions | one long prompt, and answers influence each other | independent parallel checks over the same cached context |
The probability is what makes this useful in code: you pick the threshold per decision, based on how costly a
mistake is. That's a plain if p_yes >= 0.9, not a prompt tweak.
The same engine runs three more ways, all using the LLM the agent already has instead of a separate setup:
| How | For | |
|---|---|---|
| Library | from hunch import judge, ajudge |
Python agents and scripts: in-process, no server |
| One-shot CLI | python -m hunch judge '<json>' (or - for stdin) |
shells, other languages, cron |
| MCP server | python -m hunch mcp (stdio, one judge tool) |
agents that call tools: the agent starts it on demand |
| HTTP server | python -m hunch |
a shared service; unchanged, and now also zero-config |
Settings are found in this order: arguments (judge(..., base_url=, api_key=, model=)), a hunch.toml, then the
agent's own variables. The first one set wins:
| Variables | |
|---|---|
| Endpoint | HUNCH_BACKEND_URL, ANTHROPIC_BASE_URL, OPENAI_BASE_URL, OPENAI_API_BASE, LLM_BASE_URL (a trailing /v1 is fine) |
| Key | HUNCH_BACKEND_KEY, ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, OPENAI_API_KEY, LLM_API_KEY. None at all is fine |
| Model | HUNCH_BACKEND_MODEL, ANTHROPIC_MODEL, OPENAI_MODEL, LLM_MODEL; if none is set and the endpoint serves one model, that one |
- Mode is
auto: one probe per model (remembered for the process) decides between one token and deliberate. It catches models with no thinking-off mode (GLM-5.3), but not models whose thinking-off answers are merely worse: DeepSeek-V4.1-Flash probes as one-token and failsqualifythat way. For those, setHUNCH_MODE=deliberate(andHUNCH_EFFORT=low);qualifytells you which models need it. - Every response says how it ran:
mode(resolved) andeffort(the thinking effort actually sent,nullin one-token mode). A requestedeffortthat was not applied adds anotesentry instead of being dropped silently. This is the only change to the HTTP response since v1.5: two added fields, andnoteswhen relevant. - The main model, not the small one.
ANTHROPIC_SMALL_FAST_MODELis deliberately not read. Runpython -m hunch qualifyon whatever you point it at. - It needs token logprobs. The endpoint has to be an OpenAI-compatible vLLM server (a gateway in front of one is
fine). Anthropic's API returns no logprobs and can't constrain the answer, so an agent running on Claude itself
should set
HUNCH_BACKEND_URLto a vLLM server; pointing atapi.anthropic.comis refused with that message. judge()blocks, and is safe to call from code that already runs an event loop.ajudge()is the async version. Both return the HTTP API's{"model", "results", "usage"}and raisehunch.HunchError(.code,.message).
MCP config for other clients:
{"mcpServers": {"hunch": {"command": "python", "args": ["-m", "hunch", "mcp"]}}}No client library is needed; it's one POST:
import httpx
HUNCH = "http://127.0.0.1:8791"
def judge(context, checks, model=None, images=None):
body = {"context": context, "checks": checks, "model": model}
if images:
body["images"] = images # vision models: http(s) or data: URLs, see Images below
r = httpx.post(f"{HUNCH}/v1/judge", json=body, timeout=60)
r.raise_for_status()
return r.json()["results"]
res = judge(
{"old": "The staging DB listens on port 5432.", "new": "Staging DB moved to port 6432."},
{"replaces": {"kind": "yesno",
"question": "Does `new` replace `old`'s value for the same thing?",
"yes_if": "same thing, changed value",
"no_if": "the same value restated, only more detail added, or a different thing"}},
)
p = res["replaces"]["p_yes"]
if p >= 0.9:
... # act automatically
elif p >= 0.5:
... # queue for review
else:
... # leave as isPatterns that work well:
- Fan out. Put every question you might need about one context in a single request. They run in parallel and share the cached context, so the questions your code ends up ignoring cost little.
- Choose from candidates, don't generate. Find candidates with code (regex, search, a list of IDs) and ask a
pickwhich one fits. The answer is always one of your candidates. - Decompose. Rather than one "how good is this?"
scale, ask three narrowscales and weight them in code. When your priorities change, you change a weight, not a prompt. - Gate expensive work. Use a cheap
yesno("does this passage answer the question at all?") before sending anything to a big generative model.
Hunch passes images through to the model as ordinary chat content parts, so any vision model on your vLLM server can answer checks about photos, and the answer is read off the logprobs exactly as for text:
res = judge({"headline": "Floods hit the area"},
{"shows_it": {"kind": "yesno", "question": "Does the photo show the event described in the headline?",
"yes_if": "the photo shows that kind of event",
"no_if": "a similar-looking but different scene (e.g. a wet street for a flood), "
"or a real photo of a different kind of event"}},
images=["data:image/jpeg;base64,..."])Measured. bench/images.py runs 188 freely licensed Wikimedia Commons photos, labelled by
hand, through eight questions a news or monitoring pipeline asks. Each question comes with its look-alike: a
flooded street vs a wet one (and photos of London's Flood Street), a protest vs a concert crowd or a memorial, an
active wildfire vs a red sunset or forest that has already burned, a road crash vs a traffic jam, a building on
fire vs one lit red by fireworks or a fire station, earthquake rubble vs a demolition site, a polling station vs a
queue, a press conference vs a lecture or TED talk. Scored with qualify's criteria, NVFP4 checkpoints on vLLM,
thinking off, one token per check:
| Model | "Does this photo show …?" | AUROC | ECE | "Does the photo show the event in the headline?" | AUROC | ECE | Flips |
|---|---|---|---|---|---|---|---|
| Gemma-4-31B-IT | ✅ 98.4% | 1.000 | 0.016 | ✅ 98.2% | 0.997 | 0.016 | 0% |
| Qwen3.5-397B-A17B | ✅ 97.3% | 0.999 | 0.013 | ✅ 97.4% | 0.997 | 0.055 | 0% |
| Qwen3.8-Flash-Next | ✅ 97.3% | 1.000 | 0.021 | ✅ 96.7% | 0.998 | 0.022 | 0–0.4% |
| Qwen3.6-35B-A3B (AWQ int4, edge-class GPU) | ✅ 97.9% | 1.000 | 0.072 | ✅ 98.2% | 0.999 | 0.085 | 0–0.4% |
Accuracy is at the p_yes ≥ 0.9 gate over 188 photos (photo task) and 274 checks (headline task: every photo with
its own event's headline, plus every real event photo with a clearly different event's headline, all of which the
three models rejected). For both Qwen models the definitions cost 0.4 points on the headline task (one check in
274), within the 1-point allowance.
What this shows: the 0.9 gate carries over to images unchanged. Calibration on photos is as good as or better than on text, and the few remaining misses are photos the models hedge on (a tree standing in Mekong floodwater, a car hanging off a quay), which is what a gate should send to review.
Limits: eight questions and 188 photos, chosen and labelled by us; well-known Commons photos may be in the
models' training data; no stock-photo or AI-generated-image detection was tested. Two lessons from building it:
auditing labels against the models' disagreements found one of ours wrong (a car-and-streetcar collision we had
labelled "no crash"), and headlines that name details ("thousands march demanding new elections") make most real
protest photos a correct no, so say in the check whether you mean the kind of event or this specific one.
Measure on your own images before relying on it: python bench/images.py shows how.
From the bundled benchmark (see Measured):
- Most accurate: a large instruct model (Qwen3.5-397B: 100%). Use it for decisions where errors are costly.
- Best mid-size: Qwen3.8-27B (99.6%, calibration error 0.058). It qualifies comfortably on a single GPU.
- Why not a small purpose-trained decision model? We measured one (Laya, 322M-421M, Apache-2.0). Asked in a form it can answer, its best checkpoint ranks these pairs at AUROC 0.953 — comparable to a 9B LLM, in a model 20x smaller and ~2.3x faster per call — but reaches only 81.2% at the 0.9 gate with ECE 0.301, i.e. a good ordering with probabilities you would have to calibrate yourself. Method, numbers, and two corrections we published against ourselves: docs/COMPARISON-laya.md.
- Generation beats size. Every Qwen3-generation model failed (Qwen3-14B scored below Qwen3-8B), while the newer Qwen3.5+ models pass. The small Qwen3.5-9B ranks almost perfectly (AUROC 0.996) but hedges on look-alikes, so use it only with a strict act threshold (≥ 0.97) and human review below that.
- Best calibrated per GPU: a mid-size dense model (Gemma-4-31B: 98.8%, calibration error 0.013). Its probabilities are the most trustworthy as probabilities, and it's light enough to run next to other workloads.
- Edge: a small MoE (Qwen3.6-35B-A3B, 4-bit: 99.6%). It fits on edge hardware and suits requests that ask only a few checks.
- A thinking-off mode is not automatically a good one. DeepSeek-V4.1-Flash answers in one token, but those answers
follow the order the options are listed in, and even with
debias = trueits calibration fails (ECE 0.192). Allowed to think for ~60 tokens (mode = "deliberate",effort = "low") it qualifies at 98.8% with ECE 0.013, tied with Gemma as the best calibrated model we have measured. See Models whose thinking-off mode is weaker. - Position-biased models work with
debias = true, at twice the calls. It reduces the bias by averaging it; it does not remove it. - Check the model has a real non-thinking mode. GLM-5.3 has none: its template always opens a thinking block, so
its first token is never an answer. Run it with
mode = "deliberate"and it qualifies at 97.5% (ECE 0.025); leave it in one-token mode and it scores 66.7% (see Models with no non-thinking mode). Runpython -m hunch selftestandpython bench/bench.py accuracy <model>before adopting a model. - Thinking models must have thinking switched off per request. Otherwise the one allowed token is spent on
reasoning. Hunch sends
reasoning_effort: "none"by default. Some chat templates needchat_template_kwargs = { enable_thinking = false }in the model'sextra_bodyinstead (seehunch.toml.example).
- Generating text (summaries, replies, extraction of free-form values). Use a generative model, or have code
find candidates and let Hunch
pick. - Arithmetic, counting, date comparison, exact matching. Do it in code.
- Multi-step reasoning. A check is a snap judgment over the context you give it. Split the problem, or use a reasoning model.
- Security boundaries. Context is data, but a model can still be steered by adversarial text inside it. Don't make a Hunch check the only thing standing between untrusted input and a dangerous action.
- Labels, not text. Answers map to single-token labels:
Y/N, lettersA… for options, and digits0–9for levels. Each check is sent as a chat completion withstructured_outputs: {"choice": [labels]},logprobs: true,max_tokens: 1,temperature: 0andreasoning_effort: "none". - Probabilities from logprobs. The label logprobs are exponentiated and renormalised over the allowed labels.
scale.valueis the expectation Σ i·p(i). - Big picks. vLLM returns at most
--max-logprobsalternatives (20 by default), so a pick with more than 15 options is split into groups. One call picks the group and one call per group picks within it, all in parallel: p(option) = p(group) × p(option | group). 300 options take 21 calls, 100 options take 8. - Position-bias correction (
debias = trueper model). Some models lean toward whichever answer is listed first. With debias on, yes/no checks are asked in both answer orders and small picks in both option orders, and the results are averaged. That doubles the calls, so only enable it for models that need it.
vLLM gotcha: use structured_outputs. Some vLLM versions silently ignore the legacy guided_choice
parameter, so the model answers unconstrained and nothing tells you. python -m hunch selftest checks that the
constraint is really enforced (it asks the model to write "hello" while restricted to A/B).
Hunch ships 240 fictional, labelled look-alike pairs (hunch/lookalikes.py): "does NEW replace OLD?" and
"do A and B say the same thing?". It deliberately includes restatements and same-value-different-thing traps,
two runs per model, with yes counted at p_yes ≥ 0.9. All results below use the checks in
hunch/lookalikes.py, which name the look-alike cases in no_if (e.g. "new restates the same
value, or is about a different thing"). Results on vLLM (NVFP4 checkpoints unless noted), measured 2026-09:
| Model | Verdict | Accuracy | AUROC | Brier | ECE | Max drift between runs |
|---|---|---|---|---|---|---|
| Qwen3.5-397B-A17B | ✅ Qualified | 100.0 | 1.000 | 0.017 | 0.049 | 0.150 |
| Qwen3.8-27B (bf16) | ✅ Qualified | 99.6 | 1.000 | 0.030 | 0.058 | – (0 flips) |
| Qwen3.6-35B-A3B (AWQ int4) | ✅ Qualified* | 99.6 | 1.000 | 0.055 | 0.116 | 0.263 |
| Gemma-4-31B-IT | ✅ Qualified | 98.8 | 0.991 | 0.013 | 0.013 | 0.011 |
DeepSeek-V4.1-Flash (mode = "deliberate", effort = "low") |
✅ Qualified | 98.8 | 0.991 | 0.013 | 0.013 | – (1.2% flips) |
| DeepSeek-V4.1-Flash (one token, debias on) | ❌ Not qualified | 92.9 | 0.994 | 0.133 | 0.192 | – (1.2% flips) |
| Qwen3.5-9B (bf16) | ❌ Not qualified | 90.8 | 0.996 | 0.192 | 0.266 | – (0 flips) |
| Qwen3-8B (bf16) | ❌ Not qualified | 76.7 | 0.838 | 0.238 | 0.242 | 0.169 |
| Qwen3-14B (bf16) | ❌ Not qualified | 71.7 | 0.806 | 0.286 | 0.289 | 0.195 |
GLM-5.3 (mode = "deliberate") |
✅ Qualified | 97.5 | 0.977 | 0.025 | 0.025 | – (0.8% flips) |
GLM-5.3 (mode = "one_token") |
❌ Not qualified | 66.7 | 0.734 | 0.194 | 0.065 | – (9.2% flips) |
| Qwen3-0.6B (bf16) | ❌ Not qualified | 35.8 | 0.758 | 0.611 | 0.630 | 0.024 |
Verdicts use python -m hunch qualify's default criteria (accuracy ≥ 90%, ECE ≤ 0.15, definitions cost at most 1 point,
≤ 2% flips). Qwen3.5-397B, Qwen3.8-27B, Gemma-4-31B, Qwen3.5-9B, GLM-5.3 and DeepSeek-V4.1-Flash were run through qualify itself; for the others the
same criteria are applied to their benchmark numbers. * Qwen3.6-35B-A3B passes on accuracy, calibration and stability;
its question-only run wasn't measured.
Why the others fail:
- DeepSeek-V4.1-Flash in one-token mode: calibration. Its thinking-off answers lean on the listed order of the answers; allowed to think, the same model qualifies. See Models whose thinking-off mode is weaker.
- Qwen3.5-9B: calibration. Real positives sit at p≈0.99, but the look-alike traps land at a median of 0.42–0.65, so it hedges rather than being confidently wrong.
- The Qwen3 generation: confidently wrong, and definitions make it worse.
- GLM-5.3 in one-token mode: it has no non-thinking mode, so its first token opens a scratchpad rather than
answering. In
mode = "deliberate"the same model qualifies at 97.5% with ECE 0.025. See Models with no non-thinking mode.
The same 240 pairs with only the question ("Does new replace old's value for the SAME thing?"), with no yes_if /
no_if (python bench/bench.py accuracy --vague <model>). A dash means the row came from python -m hunch qualify,
which until v1.2.1 recorded only accuracy for its question-only run; those backends are no longer up to re-measure.
| Model | Verdict | Accuracy, look-alikes named | Accuracy, question only | AUROC, question only | Brier, question only |
|---|---|---|---|---|---|
| Qwen3.5-397B-A17B | ✅ Qualified | 100.0 | 83.3 | 0.884 | 0.182 |
| Qwen3.8-27B | ✅ Qualified | 99.6 | 80.0 | – | – |
GLM-5.3 (mode = "deliberate") |
✅ Qualified | 97.5 | 81.2 | – | – |
| Gemma-4-31B-IT | ✅ Qualified | 98.8 | 80.4 | 0.863 | 0.196 |
DeepSeek-V4.1-Flash (mode = "deliberate", effort = "low") |
✅ Qualified | 98.8 | 83.3 | 0.875 | 0.167 |
| DeepSeek-V4.1-Flash (one token, debias on) | ❌ Not qualified | 92.9 | 78.3 | 0.865 | 0.241 |
| Qwen3.5-9B | ❌ Not qualified | 90.8 | 81.7 | – | – |
| Qwen3-8B | ❌ Not qualified | 76.7 | 81.2 (definitions hurt) | – | – |
| Qwen3-14B | ❌ Not qualified | 71.7 | 74.6 (definitions hurt) | – | – |
GLM-5.3 (mode = "one_token") |
❌ Not qualified | 66.7 | 67.5 (definitions hurt) | – | – |
The model and the service are the same; only the definitions changed. On every model that can answer the question at
all, two short definitions (yes_if / no_if) are worth 9–24 accuracy points. Definitions go the other way on the
older Qwen3 models, and on GLM-5.3 in one-token mode — where nothing in the prompt can help, because the token being
read is not an answer. The same GLM in deliberate mode gains 16 points from them. That contrast is exactly what
qualify's "definitions help" check is for. Write the definitions first, then choose the model.
These are results for one synthetic task on our hardware. Measure on your own labelled cases before a decision depends on
Hunch: python bench/bench.py accuracy <model> shows how.
GLM-5.3 scores 66.7% here, which looks like a weak model. It isn't — we were measuring the wrong thing.
GLM-5.3 has no non-thinking mode. Its chat template always appends a thinking block to the generation prompt, so
the first generated token is the model opening a scratchpad, never an answer. Constraining that position to Y/N
projects a distribution that isn't about the verdict onto two tokens. The result looks exactly like what we measured:
probabilities pinned near 0.5, unstable between identical runs (52 of 240 answers flipped), and very sensitive to how
the prompt is built. There is no flag that avoids this: the GLM-4.x thinking toggle was removed, and
reasoning_effort: "none" is not a recognised value — it falls through to maximum reasoning while also switching
the server's reasoning parser off.
Let the same model think, and constrain only its final verdict token, and it is excellent:
| How GLM-5.3 is asked | Accuracy | AUROC | ECE | Flips | Verdict |
|---|---|---|---|---|---|
One constrained token (mode = "one_token") |
66.7% | 0.734 | 0.065 | 9.2% | ❌ |
Thinking first, verdict token constrained (mode = "deliberate") |
97.5% | 0.977 | 0.025 | 0.8% | ✅ |
Same model, same server, same 240 pairs. It costs hundreds of thinking tokens and seconds per decision instead of
one token in well under a second — and with a 512-token budget a few checks never reach a verdict, so
think_budget = 1024 is the setting that qualified.
What to take from this:
- Hunch measures something narrow: does the model have an answer at the first generated token? A model whose template always starts with a scratchpad cannot, whatever its ability.
- The symptom is recognisable. Probabilities clustered near 0.5, unstable between identical runs and moving a lot when you reshape the prompt, mean you are reading a position that isn't the answer — not that the model is weak. Check whether your model has a real non-thinking mode before blaming it.
- A reasoning model can still give you a calibrated probability. Let it think, constrain the final answer to your
labels, and read the logprobs of that token. That is what
mode = "deliberate"does.
DeepSeek-V4.1-Flash is the opposite case to GLM. It does have a working non-thinking mode: with
reasoning_effort: "none" its first token is a clean Y or N. But that snap answer is not the model's best
judgment:
| How DeepSeek-V4.1-Flash is asked | Accuracy | AUROC | ECE | Flips | Verdict |
|---|---|---|---|---|---|
| One token | 82.9% | 0.968 | 0.225 | 2.5% | ❌ |
One token, debias = true |
92.9% | 0.994 | 0.192 | 1.2% | ❌ calibration |
mode = "deliberate", default effort |
97.9% | 0.984 | 0.021 | 0.8% | ✅ |
mode = "deliberate", effort = "low" |
98.8% | 0.991 | 0.013 | 1.2% | ✅ |
Every error in one-token mode was a false yes, and most were the restatement trap, answered with confidence:
asked whether "Port 5432 is where the staging Postgres accepts connections" replaces "The staging Postgres listens on
port 5432", it said yes at 0.99 when Y was listed first, and mostly no when N was. debias averages the two
orders, which dilutes that bias into hedged probabilities but cannot remove it — hence good ranking (AUROC 0.994) and
failed calibration. Allowed to think, it gets the same pair right in both orders, spending ~60 thinking tokens (about a
second per check on our deployment), and low effort is enough.
What to take from this: when a model can think, its thinking-off answers may be the weaker of the two, and
debias can hide that rather than fix it. python -m hunch qualify now tries both (see below).
For exactly these models, Hunch can read the answer differently. Set mode on the model in hunch.toml:
mode |
What happens | Cost per check |
|---|---|---|
one_token (default) |
The answer is the first generated token | 1 token, well under a second |
deliberate |
The model thinks; only its final verdict token is constrained, and that token's logprobs give the probability | up to think_budget tokens (2048 by default) and seconds |
auto |
One unconstrained probe at first use: if the first token opens a scratchpad, deliberate, else one_token |
one extra call, once |
Everything downstream is identical — same API, same probs and confidence, same thresholds, same qualify
criteria. In deliberate mode Hunch drops the fields that switch thinking off (reasoning_effort: "none",
enable_thinking: false, thinking: false) but keeps real effort levels.
How long it thinks: effort. On backends that support it (GLM-style reasoning_effort), set effort per model
in hunch.toml, or per request for one call. More thinking buys stability, and the curve is steep at the cheap end —
GLM-5.3 on the 240 pairs:
effort |
Thinking tokens | Accuracy | ECE | Flips between runs | Qualifies |
|---|---|---|---|---|---|
low |
~3 | 94.2% | 0.088 | 4.6% | no (stability) |
high |
~20 | 96.7% | 0.033 | 2.5% | no (just over) |
default (max) |
~79 | 97.5% | 0.025 | 0.8% | yes |
low costs barely more than one-token mode, which suits cheap high-volume checks where you re-run or review the
borderline cases; the default is what belongs behind a 0.9 gate. Measure it on your own task.
think_budget is a cap, not a knob. Raising it does not make the model think longer, and lowering it does not
make it think less — it only decides whether a long answer arrives or the call raises. Measured on GLM at default
effort: median 128 thinking tokens, p95 671, max 3579, so a 512 cap truncated 3% of checks. With effort set the
tail collapses (low max 59, high max 138) and 256 is plenty. Size it above your p99 and control cost with
effort.
python -m hunch qualify handles both situations by itself. If a model fails in one-token mode and it can think —
either because its first token always opens a scratchpad (GLM) or because it thinks once thinking isn't switched off
(DeepSeek) — it re-runs the model in deliberate mode, reports both, and tells you which setting qualifies. A model
counts as qualified if either mode passes.
- Name the look-alike. Say what counts as no for the case that looks most like yes. On the bundled
benchmark, dropping
yes_if/no_ifcosts 17–24 accuracy points on every model tested (see vague checks). One sentence inno_if("restating the same value, or a different thing") is most of the fix. Look for more than one look-alike: on memory-style pairs, a "replaces" or "duplicate" check that doesn't also say adding detail is not a replacement (or a duplicate) called refinements ("machine parts" → "machine parts, mostly pumps and valves") replacements at p = 1.0 on one model. A threshold can't catch that; naming the case fixes it. - Definitions go in
yes_if/no_ifand option descriptions, not only in the question. - One judgment per check. Split "is this a good candidate?" into several checks and combine them in code.
- Keep arithmetic, counting, dates and IDs in code. Give the model only the judgment.
- Add an escape option (
"none": "none of the above") to picks when nothing may fit. A pick always returns something. - Treat
p_yesandconfidenceas uncalibrated for your task until you've checked them against 20–40 labelled cases.
hunch.toml (path in HUNCH_CONFIG, or ./hunch.toml). See hunch.toml.example:
[backend]
url = "http://localhost:8000"
[service]
default_model = "qwen"
max_concurrency = 16 # simultaneous backend calls
# api_keys = ["change-me"] # required when listening beyond localhost
[models.qwen]
backend_model = "Qwen/Qwen3.5-397B-A17B"
[models.deepseek]
backend_model = "deepseek-ai/DeepSeek-V4.1-Flash"
mode = "deliberate" # its thinking-off answers are position-biased; thinking briefly fixes that
effort = "low"| Environment variable | Meaning |
|---|---|
HUNCH_HOST / HUNCH_PORT |
Listen address (default 127.0.0.1:8791) |
HUNCH_BACKEND_URL, HUNCH_BACKEND_KEY |
vLLM server and optional bearer key |
HUNCH_BACKEND_MODEL |
Quick single-model setup without a TOML file (exposed as default) |
HUNCH_DEFAULT_MODEL |
Overrides service.default_model |
HUNCH_CONCURRENCY |
Simultaneous backend calls. Keep it modest on a shared server |
HUNCH_BACKEND_MODEL + mode in hunch.toml |
one_token (default), deliberate or auto — see Deliberate mode |
SSL_CERT_FILE |
CA bundle for a backend behind a private or self-signed CA, e.g. /etc/ssl/certs/ca-certificates.crt (Python doesn't use the system store by default) |
HUNCH_API_KEYS |
Comma-separated bearer keys. Hunch refuses to listen beyond localhost without them (or HUNCH_ALLOW_NOAUTH=1 behind an authenticating proxy) |
- Docker:
docker build -t hunch ., then run it withHUNCH_BACKEND_URL,HUNCH_API_KEYSand a mountedhunch.toml(seedeploy/docker-compose.yml). The image runs as non-root undertini, with a health check. - systemd:
deploy/hunch.service. - After deploying, run
python -m hunch selftest(ordocker compose exec hunch python -m hunch selftest).
SKILL.md is a drop-in guide for an agent writing a Hunch integration: the call, the rules for writing
checks, thresholds, qualifying a model, and what not to do. llms.txt indexes it for agents that look for
one. Point your agent at the raw URLs:
https://raw.githubusercontent.com/ihubanov/hunch/main/llms.txt
https://raw.githubusercontent.com/ihubanov/hunch/main/SKILL.md
pip install -e '.[dev]'
pytest -q # offline tests against a fake backend
python -m hunch selftest # quick live pre-flight against your backend
python -m hunch qualify qwen # full qualification of a model (verdict + reasons)
python bench/bench.py accuracy qwen gemma # live accuracy benchmark
python bench/bench.py fanout qwen # latency vs number of checksMIT. See LICENSE.