Independent implementation of the documented System One wire format — not affiliated with TypeSafe.
method= picks how the model is asked to decide, and how its answer is turned back into a distribution. The
default, auto, selects either logprobs or structured per client, model and surface; grammar and
discrete are explicit choices (see auto). All four methods share the same
label machinery: options are labelled A, B, C, … in criteria order, and label_to_key maps a label back
to the option key (Choice), the zero-based level index (Score) or True/False (Noul). Switching
methods never changes your question or answer types — only the request body and the readout.
For how a method fits into the layers around it — prompt, transport, surface — see Architecture.
Labels are single letters while a question has 26 options or fewer. Past that they become two letters (AA,
AB, … ZZ), which only structured and discrete can use: they answer in JSON, where a label is just a
string. logprobs and grammar read the label token, and the first token of "AA" is "A", so they raise
InvalidQuestionError past 26 options and name the two methods that can take more (up to the Jev API limit of
255).
With a single option there is no distribution to read: the sampled label is the answer, so logprobs
and grammar report it with probability 1.0 and confidence 1.0 whether or not the provider sends
alternatives. The two-candidate rule below applies to a question that has something to compete with.
| Method | Asks for | Distribution comes from |
|---|---|---|
logprobs |
one label, plus the logprobs of the alternatives | the model's own next-token distribution |
grammar |
one label, constrained by a GBNF grammar | the model's own next-token distribution |
structured |
a JSON object with a probability per option | the model's stated numbers, validated and normalized |
discrete |
a single option, as JSON | one-hot over the chosen option |
Pick logprobs when the provider returns chat logprobs: it is one short call, and the numbers are the model's
real distribution rather than a self-report. Pick structured when logprobs are unavailable or the provider
supports strict JSON schema, and discrete when you only need the decision and want to skip probabilities.
grammar exists for self-hosted Chat Completions servers that accept a grammar field. Leaving method at
auto normally adds one answer request for a failed logprob attempt before discovery, but transient retries
and field downgrades can add more; two-step reasoning can add two calls per pass (see auto below).
method="auto" — the default — answers with logprobs where the provider returns them and with structured
where it does not, so the same code works against a logprob-capable server, a reasoning model, and a provider
that never implemented logprobs. The choice is made per (model, surface) by observation and remembered for the
life of the client.
Four kinds of evidence count as this provider cannot do logprobs:
| Evidence | Response |
|---|---|
The provider rejects the logprob fields with HTTP 400, 403 or 422 and names them — Gemini's OpenAI-compatibility layer answers Unknown name "logprobs": Cannot find field., while a reasoning model behind an OpenAI-shaped gateway answers logprobs are not supported with reasoning models. |
Under api="auto", ask on the surface that carries the readout — and remember the verdict, so later calls start there |
The provider refuses the include entry a Responses request carries them in, without ever writing the word "logprob" — OpenRouter answers 400 Invalid option: expected one of … for path: ["include", 0], and OpenAI's own wording for a model that offers no includable is 400 Unsupported parameter: 'include' is not supported with this model. |
same. The carrier is per surface: on Chat Completions it is the logprobs field, so a message that merely mentions include is about something else |
The answer carries no logprobs at all (logprobs: null, or a compatibility layer that drops the field) |
same, once a second response confirms it |
The answer token's logprobs carry no alternatives — top_logprobs empty, or nothing but the sampled token — so there is no distribution to read |
same |
With method="auto", a label readout that has no other usable surface — a client that speaks only one, a route
already known to be missing, or reasoning="native" — answers that question with structured. A provider
refusal is remembered immediately; a weak readout absence is remembered after the second one. grammar never
enters this fallback because it is explicit. Pinned logprobs may move to another surface under api="auto",
but it is never swapped for structured and reports the provider failure when no surface is left. Pinned
grammar is never substituted; surface selection and the grammar check determine whether its request is sent.
Under method="auto", a 5xx that survives the transient retries answers that question with structured, but
is not remembered as a capability verdict, so the next question tries logprobs again. With a pinned logprobs
method, the same failure reaches the caller as a provider error.
Only the last two table rows are weak evidence: they describe a response whose logprobs could not be read. A
truncated answer is classified before readout and raises IncompleteAnswerError. A reasoning-only response
that is not classified as truncated raises LabelReadoutError after its configured corrective retry; neither is
treated as absent logprobs. auto answers a weak absence with structured right away but only stops asking
for logprobs after a second such response. A readable distribution in between resets the count because it
proves the provider can do it.
Cost and consequences:
- At the default
max_concurrency=8, up to four questions can each make a logprob attempt if they are already in flight when discovery starts; the code does not guarantee exactly four.max_concurrency=1makes exactly one logprob attempt because later questions take the verdict. After a field refusal or the second weak absence, laterautocalls start with the resolved method; the first weak absence answers withstructuredbut is not cached. On one surface, a failed two-step label attempt followed by the structured fallback uses four successful responses — analysis, label answer, analysis, structured answer — before allowing for retries and downgrades. Moving surfaces can change that count because it recomputes the reasoning mode. - The fallback asks for the model's own probabilities, which are a different quantity from a token
distribution.
debug["methods"]says which method each question used, and the successful attempt'sdebug["llm_attempts"][-1]["readout"]["source"]says which readout produced its answer. Failed attempts have no readout source. - A
Choicewith more than 26 options is answered in JSON without ever asking for logprobs: one label token cannot distinguishAAfromA. - The verdict lives on the client instance and is keyed by model and surface: a new client, or an explicit
method="logprobs", starts over. Passingmethodtosystem_oneoverrides it for that call. - Only a provider's field-refusal verdict is remembered at once; a response-level absence is remembered on the
second one (see above). Under
method="auto", a 400, 403 or 422 that complains about the value sent — for example,Invalid 'top_logprobs': integer must be between 0 and 5, but got 20.— answers that question withstructuredwithout caching anything, so the next call tries logprobs again. With pinnedlogprobs, the provider error reaches the caller instead.
Provider support, as of this release — check your provider's docs, since this moves:
| Provider | logprobs |
Note |
|---|---|---|
OpenAI gpt-4o, gpt-4.1 |
yes | |
OpenAI reasoning models (o-series, gpt-5 family) |
no | 400 logprobs are not supported with reasoning models. |
| OpenAI Responses surface | partial | include alone returns the sampled token and no alternatives; some models fail outright on top_logprobs >= 2 |
| Anthropic Claude | no | no logprob API at all |
| Gemini via the OpenAI-compatibility endpoint | no | 400 Unknown name "logprobs": Cannot find field. |
| Gemini native API | yes | not reachable through an OpenAI-compatible client |
| DeepSeek | yes | top_logprobs up to 20 |
| Together | yes | send top_logprobs for alternatives; logprobs: 1 alone returns the sampled token |
| llama.cpp | yes | Chat Completions only: its /v1/responses shim rejects the logprob fields (400 top_logprobs requires logprobs to be set to true), so auto re-asks on Chat Completions |
| vLLM | yes | caps top_logprobs at its own --max-logprobs (20 by default); /v1/responses carries them through include |
| SGLang | yes | its /v1/responses needs top_logprobs sent explicitly (it defaults to 0) — jevper always sends it |
| Ollama | partial | local builds since Nov 2025 return logprobs on Chat Completions; its /v1/responses returns an empty logprob list, so auto re-asks on Chat Completions. Ollama Cloud and older builds report none at all |
| OpenRouter | per model | it routes by price and, by default, sends your request to an endpoint that may ignore logprobs — the answer comes back with none, which auto reads as "no logprobs" and falls back on. Add extra_body={"provider": {"require_parameters": True}} to route only to endpoints that support every field you send. Its Responses API rejects the logprob includable outright (400 Invalid option: expected one of … at path: ["include", 0]), so an auto readout moves to Chat Completions, where the distribution arrives — verified live, with an explicit method="logprobs" as much as with auto. It also takes a prompt_cache_key longer than OpenAI's 64 characters without complaint, and answers an unknown model id with 400 is not a valid model ID rather than a 404 |
| everything else | unknown | reasoning models and thin compatibility layers are the ones that say no |
With auto you do not have to know this table.
flowchart TD
A["select_surface(client, api, method)"] --> B{"api"}
B -->|"responses"| C["require client.responses.create"]
B -->|"chat_completions"| D["require client.chat.completions.create"]
B -->|"messages"| G["require client.messages.create"]
B -->|"auto"| E{"method == grammar"}
E -->|"yes"| D
E -->|"no"| F{"client has responses.create"}
F -->|"yes"| C
F -->|"no"| H{"client has chat.completions.create"}
H -->|"yes"| D
H -->|"no"| G
C -->|"404 not identified as a model error"| D
api="auto" (the default) prefers the Responses surface because it carries native reasoning and encrypted
content, except for grammar, which only Chat Completions can carry. messages — the Anthropic-compatible
API — is the last choice. Its normalized responses carry no logprobs, so a client whose only surface is
messages answers auto with structured; explicit logprobs raises UnsupportedMethodError before a
request. An explicit grammar request with api="messages" raises UnsupportedMethodError after that surface
is selected. With api="auto", a messages-only client has no Chat Completions route, so grammar instead
raises ClientCapabilityError during surface selection. A missing attribute likewise raises
ClientCapabilityError naming the surface to pass explicitly.
A client object cannot tell you whether the server implements the route: openai.OpenAI exposes
responses.create either way, so a server that does not implement it answers 404 for that call. Under auto,
that 404 is read as a missing route unless the error identifies a model error: it either has a recognized
model-error code, or names the requested model together with a model-404 marker. An unknown-model code without
the model id therefore does not switch surfaces. A route 404 is re-issued on chat_completions and remembered
for the client's life. An explicit api="responses" is a decision, not a preference: its 404 reaches you
unchanged.
The remembered verdict only ever skips a route, so it moves the call only when the client can speak the
other surface. A client whose only surface is messages stays on it, pays the 404 again, and reports it —
the same error the call that learned the verdict raised, rather than an AttributeError for an attribute the
client never had.
The preference has one exception, and it is about the readout rather than the surface. A server can implement
the Responses route and still not carry logprobs through it: ollama answers it with an empty logprob list,
llama.cpp refuses the logprob fields there outright (400 top_logprobs requires logprobs to be set to true),
and OpenAI's Responses logprobs hold the sampled token with no alternatives. When the label readout cannot
produce a distribution on the chosen surface — for any reason other than a provider failure that survived its
retries — auto re-asks on the other surface, marks the one that failed so later calls start where the
distribution is, and keeps that mark for the model. Three things stop the move: reasoning="native", because
native reasoning is the reason to prefer Responses and switching would silently turn it into a two-step pass;
a grammar request, which is a Chat Completions convention the other surface cannot carry; and a server whose
other route is missing or already known to be missing. A distribution arriving later on a
marked surface clears the mark: the verdict moves the readout, it does not condemn the surface.
For the four servers this was checked against — what to pass, how to turn thinking off, and what fits a 12 GB card — see local-servers.md.
Request fields per surface:
| Chat Completions | Responses | |
|---|---|---|
| messages | messages=[...] |
input=[...], plus store=false |
| logprobs | logprobs=true, top_logprobs=N |
top_logprobs=N, include=["message.output_text.logprobs"] |
| JSON schema | response_format={"type": "json_schema", "json_schema": {"name": ..., "schema": ..., "strict": true}} |
text={"format": {"type": "json_schema", "name": ..., "schema": ..., "strict": true}} |
schema fallback (structured_outputs=False) |
response_format={"type": "json_object"} |
text={"format": {"type": "json_object"}} |
| grammar | extra_body={"grammar": "..."} |
not available |
| reasoning | reasoning_effort (only when effort is set) |
reasoning with whichever of effort, summary, and context are set |
The Chat Completions and Responses builders do not set output-token caps. Label readouts inspect the first
answer token, while structured and discrete parse the complete JSON answer. The Messages builder always
sets the required max_tokens. On the two OpenAI surfaces, any other provider field goes through
extra_body.
The Messages surface has no logprobs at all, and it does have a schema field of its own: Anthropic's
output_config={"format": {"type": "json_schema", "schema": ...}}, the counterpart of the two above and
what the TypeSafe reference adapter sends there. jevper sends it and keeps the schema in the system
prompt: vLLM implements the field (a schema naming a constant the prompt never mentions comes back with
that constant in the answer), while llama.cpp and LM Studio accept it and ignore it, which no error reports
— and ollama and SGLang accept it too, though with a thinking model nothing comes back on that route to
enforce it. A server that discards a field it accepted looks exactly like one that never read it.
The schema is also rewritten for Anthropic's documented subset on the way out (numerical constraints are a
400 there), and the whole field travels in the request body rather than as an SDK keyword, since the
oldest Anthropic SDK jevper supports has no such parameter.
When a server refuses an optional request field, jevper drops it and re-asks, then remembers the limit for the
rest of the client's life. The order is surface-specific. Responses evaluates the include list first: a
refusal of reasoning.encrypted_content drops that entry before the schema and reasoning rungs are considered.
Chat Completions and Responses use json_schema → json_object → no format field. Messages has no
json_object rung, so refusing output_config drops directly to the prompt-only schema. After the schema rung,
an unrecognized reasoning field, prompt-cache key or thinking field can be dropped, except when the complaint
is about the value sent. The prompt already asks for one JSON object and the readout validates it, so a
question is answered instead of failing. The ladder is finite, so a server that refuses every optional field
ends in a ProviderError, and debug["server_limits"] reports what was learned. A caller-owned copy in
extra_body is removed with the learned limit because the SDK merges extra_body last.
The include list carries two entries on a Responses call that asks for both a label readout and reasoning.
For include[1]: expected one of "message.output_text.logprobs", jevper reads the refusal as applying to
reasoning.encrypted_content, disables that entry and re-asks with only the logprob carrier, keeping the
method. It does not omit the whole include field. The accepted values printed by a server are read as the
list it supports; the word "logprob" in that list is how the server names what it will carry.
Two extra_body fields interact with jevper's own rather than replacing it. logprobs and top_logprobs are
separate: naming only a truthy logprobs still gets jevper's alternatives when top_logprobs is absent.
With only logprobs: false, jevper omits its typed top_logprobs, but a caller-supplied top_logprobs key
remains authoritative and is still sent.
On Messages, an enabled caller thinking object with an integer budget_tokens gets the same local rule as
ReasoningConfig(budget_tokens=…): the budget must be strictly below max_tokens, so jevper raises when a
caller-supplied cap is too small. A disabled thinking object has no budget, and a caller-supplied
extra_body["temperature"] remains authoritative. When jevper supplies the thinking block, it currently sends
{"type": "enabled", "budget_tokens": N}. Anthropic now marks manual budgets deprecated on Claude 4.6 and
rejects them on 4.7+, where {"type": "adaptive"} with output_config.effort is supported; the 1024 minimum
and strict max_tokens ceiling still apply where manual mode is available.
A refusal of a value is not automatically a refusal of the field. budget_tokens: must be at least 1024,
Invalid 'top_logprobs': integer must be between 0 and 5, and reasoning_effort must be one of low, medium, high identify fields the server knows, so reasoning, cache-key and thinking downgrades are suppressed and the
provider error travels back without a learned limit. Schema refusals are different: a complaint that reaches
the schema markers advances the schema ladder even when it concerns the schema's contents. Other capability
fields advance only when their field markers and the relevant capability evidence match.
Request: the label prompt plus logprobs=true and top_logprobs (default 20). The constructor accepts 0, but
pinned logprobs and grammar require at least 2 and reject smaller values before a request. auto can send
0; if the provider returns no usable alternatives, it falls back to structured. Responses carries the
readout in include=["message.output_text.logprobs"].
Readout:
- The first non-whitespace token of the answer must be a label, compared case-insensitively over ASCII only:
ais labelA, andı(whichstr.upperwould fold toI) is not a label at all. A reasoning server reports logprobs for every generated token — vLLM, SGLang and ollama include the thinking span whilemessage.contentholds only the answer — so the answer's own tokens are located first, by matching the answer text against the tail of the token stream. The match is strict: without a separated trace, or without an exact tail, nothing is skipped and the stream is read as it arrives, so a mismatch is an error rather than a guess. Disabling thinking is still the better deployment — a one-token answer does not need a reasoning pass. - Its logprob is taken from the token itself, and the rest of the distribution from the token's
top_logprobsentries. A label the provider did not report gets probability exactly0.0and is listed indebug["labels_missing"]. A provider that reportslogprob: null— some OpenAI-compatible servers do — is treated the same way: no number is invented, so a missing alternative is0.0and a missing logprob for the answer token raisesLabelReadoutErrorinstead of reading as certainty. - The logprobs are softmaxed over the labels only. A grammar masks logits but never renormalizes them, so
renormalizing over the label set gives the post-mask distribution — which is why
grammarreuses this readout unchanged.
For criteria {"billing", "technical", "sales"} and logprobs A:-0.12, B:-2.47, C:-3.48, the answer is
billing with {"billing": 0.884873983, "technical": 0.084389690, "sales": 0.030736327} and
confidence = 0.827310974.
Caveats:
top_logprobs=0reports nothing but the answer token, and one logprob is not a distribution: the readout raises rather than reporting certainty, andautofalls back tostructuredinstead. Raisetop_logprobs(20 covers 21 options) to get a real distribution.- A truncated
top_logprobslist silently zeroes the missing options; checkdebug["labels_missing"]when that matters. - The distribution is the model's preference over the next token, so the prompt must leave the label as the
only sensible continuation — that is what the system prompt and the
Options:block are for. - OpenAI reports
-9999.0for tokens outside the top 20 rather than omitting them; that underflows to0.0like any other very low logprob, so it needs no special handling. - Option descriptions and instructions are rendered verbatim into the options block. A description containing a
newline followed by a label-shaped line (
B: something) injects a pseudo-option into the prompt; the readout only accepts the allocated labels, so the result is aLabelReadoutErrorand one corrective retry, but keep descriptions single-line and free of label-like lines.
Failure modes, all raising LabelReadoutError or a subclass:
- No distribution at all: no logprobs came back (
no logprobs returned for the answer token (method='logprobs')), or the answer token'stop_logprobsheld nothing but the sampled token. The message namesstructuredanddiscrete, and no corrective retry is spent — re-asking with a correction turn cannot make a provider report logprobs it does not have.autoanswers these withstructuredinstead of raising. - Unusable answer: no non-whitespace token, or a first token that is not a label (
first non-whitespace token 'The' is not one of the labels [...]). These are retried once with a correction message: the model, not the provider, is at fault. - Unusable number: no logprob for the answer token, or a non-finite logprob (
nan/inf) from the provider — anandistribution would otherwise poisonconfidenceandscore.
Request: identical to logprobs, plus a GBNF grammar for the labels, e.g. for three options:
root ::= "A" | "B" | "C"
The grammar is merged into extra_body as {"grammar": "..."}, the field llama.cpp's OpenAI-compatible server
reads on /v1/chat/completions. Readout is the logprobs readout, unchanged.
Constraints:
- Chat Completions only. With
api="responses"orapi="messages", the selected non-Chat surface reaches the grammar check and raisesUnsupportedMethodErrorbefore a request. Withapi="auto", surface selection requireschat.completions.create; a messages-only client therefore raisesClientCapabilityErrorinstead. - The server must still return logprobs; if it does not, the readout raises
LabelReadoutErrorsuggestingmethod="discrete", which skips probabilities, and no corrective retry is spent — another turn cannot change what the provider reports. - Most hosted providers reject or ignore an unknown
grammarfield, so this is a self-hosted-server method.
Request: a strict JSON schema, with the model reporting a probability per option. Schema names are
jevper_choice, jevper_noul and jevper_score; every object sets additionalProperties: false and lists
all properties in required. Each probability is bounded (minimum: 0, and for noul also
maximum: 1): strict mode accepts those on the OpenAI surfaces, and the constraint that matters here — the
distribution summing to 1 — is not expressible in JSON Schema, so that check is client-side either way. The
Messages surface is the exception: Anthropic's structured outputs reject numerical constraints outright, so
each bound is folded into the description of the field it bounded before the request goes out, and the
prompt still carries the full schema.
{"type": "object",
"properties": {"probabilities": {"type": "object",
"properties": {"billing": {"type": "number", "minimum": 0},
"technical": {"type": "number", "minimum": 0},
"sales": {"type": "number", "minimum": 0}},
"required": ["billing", "technical", "sales"], "additionalProperties": false}},
"required": ["probabilities"], "additionalProperties": false}noul uses {"noul": {"type": "number", "minimum": 0, "maximum": 1}} (the probability of true); score
uses the level indexes "0", "1", … as keys.
Readout:
- Parse the answer with
json.loads. If that fails, scan each{-delimited candidate in order until one decodes, so prose braces and fences can be skipped; a second decodable object is rejected as contradictory. - The object must carry exactly the one field this method asks for —
probabilities, ornoulfor aNoulquestion.choicethen requires exactly the option keys, each a finite number>= 0;noulrequiresnoulin[0, 1]and expands to{True: v, False: 1 − v};scorerequires exactly the level keys. A missing or extra key, a wrong key set, booleans,NaNand negatives are malformed — an object that answers twice ({"probabilities": …, "choice": "technical"}) is a model contradicting itself, and reading the field jevper happened to pick would report one of the two as the answer. Key-set errors include bounded key names; numeric-validation errors also include the invalid value. This keeps a body padded with provider text from copying an unbounded payload into every error and retry reason. - Normalization (not part of the readout): when
abs(sum − 1) > 1e-6andnormalize_probabilities=True(the default), the distribution is rescaled to sum 1 and both the error and the model's original numbers are recorded indebug["probability_errors"]anddebug["original_probabilities"]. A zero total becomes uniform. Values whose sum leaves the float range — a model that answers1e308three times — are scaled by their largest value first, so normalization cannot raiseOverflowError. Withnormalize_probabilities=Falsethe model's numbers are returned verbatim and only the error is recorded — never raised, matching the reference adapter. choiceis the argmax of the final distribution, andconfidenceis computed from it. AscoreisΣ i·pᵢover the distribution rescaled to sum 1, whateverprobabilitiesreports: an expected value read off an unnormalized distribution leaves the0..N-1line the Jev answer schema documents, and this is the arithmetic the reference adapter uses.
structured_outputs=False keeps the schema in the prompt but sends {"type": "json_object"} on Chat
Completions and Responses. With strict outputs enabled, a server refusing the strict schema is re-asked with
json_object; a server refusing that too is answered with the prompt-only schema. Messages has no
json_object rung: refusing strict output_config drops directly to the prompt-only schema. The answer is
validated client-side, so an unusable shape raises MalformedAnswerError after the corrective retry.
temperature=0.0 is worth setting here (and for discrete): the answer is a single sampled JSON object, so
sampling noise moves the probabilities directly.
Request: a strict JSON schema asking for one option and nothing else.
| Question | Schema (jevper_choice / jevper_noul / jevper_score) |
|---|---|
choice |
{"choice": {"type": "string", "enum": ["A", "B", "C"]}} |
noul |
{"noul": {"type": "boolean"}} |
score |
{"score": {"type": "integer", "enum": [0, 1, 2]}} |
Readout: the object must carry exactly choice, noul or score, one-hot over the chosen option. choice
accepts a label ("B", case-insensitively over ASCII) or a criteria key ("billing"); an exact criteria key
wins over a label spelled the same way, so an option keyed "a" is read as that option and not as the first
label, and a key that is itself spelled with surrounding spaces is that key rather than a stripped version of
it. noul requires a JSON boolean or its string form ("true"/"false"); score requires a level index — an
integer, an integral float such as 2.0, or a number in a string such as "2" or "2.000", decided on the
string rather than on the float it converts to, so "2.0000000000000000000001" is malformed rather than
rounded to level 2. The string forms are what models tend to emit when the schema is only in the prompt
(structured_outputs=False). A bool is rejected as a score, since true would otherwise read as level 1.
Anything else raises MalformedAnswerError.
Because the readout parses a JSON object, the system prompt, the few-shot demonstrations and the two-step answer cue all ask for that JSON object rather than for a bare label.
The resulting confidence is 1.0 for choice and score — all the mass sits on one option, which is
maximal confidence under both formulas. noul answers carry no confidence.
MalformedAnswerError and model-answer LabelReadoutError are recoverable: the client appends a correction
turn and re-issues the answer call up to n_retry_malformed times (default 1). The correction asks for exactly
one of the labels or for only a JSON object matching the schema. Reasons are recorded in
debug["retry_reasons"]. usage.n_calls counts every successful provider response, including a response that
later fails readout and a successful corrective retry; a failed provider attempt appears in
debug["llm_attempts"] but does not increment n_calls.
A provider-unavailable LabelReadoutError — no logprobs at all or no alternatives for the answer token — is not
corrected, because another turn cannot change what the provider returns. method="auto" answers those
questions with structured; pinned logprobs or grammar reports the error.
When there is no answer to read at all, the message says why, because the parse error alone sends a caller looking for a bug that is not there:
- The budget ran out. Each surface has its own word for it — Chat Completions
finish_reason: "length", the Messages APIstop_reason: "max_tokens", the Responses surfacestatus: "incomplete"withincomplete_details.reason: "max_output_tokens"(or the"max_tokens"spelling OpenAI's own streaming example uses) — and all of them reach the caller as the provider ran out of output tokens before the answer was complete, with the field to raise on that surface:max_output_tokensfor Responses,max_completion_tokensfor Chat Completions (withmax_tokensnamed as the local-server alternative), andmax_tokensfor Messages. This holds however the answer was cut off: mid-object, or with no{at all. The Messages API's other reason,model_context_window_exceeded, is the same failure with the opposite remedy — the request is already too long to answer in — so it arrives as the provider's context window ran out, naming the state and the examples as what to shorten. A stop reason that is not one the surface documents — or not a string at all — is reported the same way rather than read. - The model refused, or the content was filtered. Each surface puts a refusal in its own place, and
jevper reads all of them: a
refusalsibling of a nullcontenton Chat Completions, arefusalcontent part on Responses,stop_reason: "refusal"on the Messages API, and a safety filter asfinish_reason: "content_filter"or a Responsesincomplete_details.reasonof the same name. A filter is a refusal in everything but the word — the content was withheld on purpose — so both arrive asModelRefusalErrorwith the provider's own report in the message, and neither is corrected: another turn is refused the same way. The model's own words ride along when it gave them, so a refusal reads as a refusal rather than as malformed JSON, and one that arrives where the answer would have been is never parsed as the answer. - The server separated reasoning from the answer and sent no answer. A reasoning parser with thinking on
does this (see
local-servers.md), and the message says so instead of "no non-whitespace token in the response".