You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Seen by the owner on the testing box (2026-09-13). As of main @ d201052.
Chat transcript (web shell): the owner writes "yeah, move them over there"; the assistant
replies with a single line — "Let me verify the current state first — the tree, any pending
suggestions, and the notes" — and the turn ends there. No further text, no tool calls, no
error banner. Just Copy / Regenerate under the partial reply.
Core-app log at the same moment (WARN, all core-app):
12:37:27.026 WARN recall skipped: backend error
{"error": "Unexpected Response: 400 (Bad Request)\n...Vector dimension error: expected dim: 768, got 4096",
"error_type": "UnexpectedResponse", "elapsed_s": 0.61}
12:37:27.252 WARN streaming turn failed
{"error": "litellm.BadRequestError: OpenrouterException - {\"error\":{\"message\":\"qwen/qwen3-embedding-8b
is an embedding model and cannot be used with the chat/completions endpoint. Use the /embeddings endpoint
instead.\",\"code\":400}}...\" LiteLLM Retried: 2 times"}
12:37:38.512 WARN recall skipped: backend error (same 768-vs-4096 payload, elapsed_s 0.21)
12:38:12.441 WARN recall skipped: backend error (same, elapsed_s 0.27)
The owner had recently switched the tenant's embedding model to a hosted OpenRouter one
(qwen/qwen3-embedding-8b, 4096 dims) through the gateway work of #868/#865. Three distinct
defects follow, filed together because they are one incident chain: a mis-scoped model picker
let an embedding model become the chat default, recall then silently lost memory instead of
healing, and the resulting mid-turn failure left the shell with a bare partial reply. Related: #879 (search doesn't self-heal a dimension mismatch the way indexing does — the knowledge/notes
half of this same class of bug); #936 and #937 are the two other recent dogfood issues on this
milestone, filed separately (websearch degrade-detection and image-PDF OCR) — siblings, not
duplicates.
Where it lives
1 — the chat/completions call went to the embedding model
The saved-hosted-models list is capability-blind end to end, and nothing downstream ever checks
a model's chat-vs-embedding mode before handing it to litellm.acompletion.
services/web/src/screens/ChatScreen.tsx:699-716 — the chat ModelPicker's free-text hosted
box calls chooseHosted(custom.trim()) on submit, which both sets the model and persists it
(ChatScreen.tsx:642-645, save.mutate(id) → api.addSavedModel).
services/core-app/src/epicurus_core_app/llm/routes.py:448-465 (add_saved_model) validates
only is_hosted(model) (llm/providers.py:64-76 — true for any <provider>/<name> string with
a known alias); it never asks whether the id is chat-capable.
services/web/src/screens/ModelsScreen.tsx:1734-1820 (SavedHostedModels) lists every
saved id under a star button wired to setDefault.mutate → api.setGlobalDefault → PUT /platform/v1/llm/prefs/default.
services/core-app/src/epicurus_core_app/llm/routes.py:312-318 (set_default) writes
whatever string it's given straight into llm_prefs.global_default — no validation at all.
This rules out the "chat and embed prefs share a key" theory: llm/prefs.py:35-37 and :105-129 keep global_default and embed_default as two distinct nullable columns; they
never collide.
services/core-app/src/epicurus_core_app/llm/gateway.py:197-203 (effective_default) returns
the stored value unchecked, and chat()/stream()/stream_chat()
(gateway.py:463-521, 522-644) hand the resolved candidate straight to litellm.acompletion(...) with no pre-flight capability check — OpenRouter is the only thing
in this call graph that ever rejects an embedding model for chat use.
The capability data needed to prevent this doesn't exist: gateway.py:977-1065
(show/_hosted_details) never reads LiteLLM's model-info mode field (chat vs embedding vs others in its cost map); it derives capabilities only from vision support and
always assumes tools for a hosted id. For an id outside LiteLLM's static map — likely for a
newer OpenRouter alias like qwen/qwen3-embedding-8b — _note_unmapped
(gateway.py:909-922) logs once and the call defaults to ["tools"]: capability-wise
indistinguishable from a genuine chat model. SavedModel.capabilities
(llm/routes.py:136-137) round-trips exactly this value to the web, so the star-picker's
badges can't tell the two apart either.
The only place this is even documented as a hazard is a sentence of help text on the embedding side: services/web/src/screens/ModelsScreen.tsx:1490-1493 — "pick an embedding model there; a chat model will fail at embed time." Nothing warns the other
direction, and neither warning is backed by real capability data.
2 — recall silently degraded after the dimension change
agent/agent.py:1483-1515 (_recall_within_budget) time-boxes the recall embed and swallows
any exception as log.warning("recall skipped: backend error", ...) (line 1510) — the turn
proceeds with no recalled facts, invisible to the owner.
This is not an uncontracted store, contrary to what the raw log suggests: the core's own
memory-fact store already carries a dimension-drift self-heal, older than #868 — memory/facts.py:94-128 (UserFactStore._ensure, #436/ADR-0074), documented at docs/services/core-app.md:759-762 and :1890-1897 as reconciling "on first use each process
lifetime". Both save() (facts.py:224) and search() (facts.py:282, which Memory.recall → UserFactStore.recall → search all funnel through, memory/memory.py:110-112)
call _ensure before touching Qdrant.
On paper the very first recall after the model switch should have reconciled the <tenant>__facts
collection (scope_collection("facts", tenant), epicurus_core/tenancy.py:81-83) to 4096-d and
every later call should succeed. The log shows the opposite: the identical raw UnexpectedResponse ("expected dim: 768, got 4096") recurs three times over ~45s
(12:37:27 / :38 / 12:38:12), each with a fast elapsed_s (0.61 / 0.21 / 0.27) — far too quick to
be an in-flight re-embed of stored facts, and inconsistent with the reconcile ever actually
running. Two candidate mechanisms in _ensure itself would produce exactly this symptom:
facts.py:117 — current_dim = vectors_config.size if isinstance(vectors_config, VectorParams) else None. If this ever evaluates None (a qdrant-client response shape the check doesn't
recognize), the drift branch is skipped entirely — no reconcile happens.
facts.py:128 — self._ensured.add(collection) runs unconditionally after the if block,
including the no-op path above. That permanently caches the collection as "reconciled" for the
rest of the process's life even though nothing was fixed, so every later search()/save()
repeats the identical fast 400 forever — matching the log pattern exactly.
Either way, the operator has no way to see "the recall store still needs a rebuild" versus "Qdrant
is unreachable" — agent.py:1508-1515 logs only the raw str(exc) / error_type at WARN, and
nothing surfaces on the Models page (contrast the knowledge/notes dimension-change contract,
ADR-0132, which is drop-and-recreate rather than in-place re-embed and evidently holds up better
under a large dimension jump).
3 — a failed streaming turn leaves the UI with a partial reply and no error
agent/agent.py:1160-1186 is the mid-stream exception handler wrapping the whole per-round tool
loop (agent.py:911-1159), so a litellm.BadRequestError raised by any round's gateway.stream_chat call (agent.py:927-928) lands here.
When nothing was produced before the failure (not (partial.strip() or timeline), agent.py:1164), a real terminal event is emitted: yield AgentEvent(type="error", detail=banner) (agent.py:1167) — and the web already renders this correctly
(services/web/src/stores/chat.ts:416-419 sets {error: detail, ...}).
When there is partial content or tool activity — the owner's exact scenario, a lead-in
sentence before/around a tool call — the code takes the other branch (agent.py:1169-1186). banner (the real reason: the raw litellm/OpenRouter message) is computed at agent.py:1162
but is never used on this branch. Only a hardcoded generic note — _STREAM_INTERRUPTED_MESSAGE = "The answer was interrupted before it finished. Please try
again." (agent.py:251) — gets appended as one more delta (agent.py:1174), followed by done (agent.py:1186). No error-typed event is ever emitted on this path, so chat.ts's working error banner is unreachable for exactly the failure mode this bug reports. _stream_failure_messages (agent.py:254-265) only special-cases a short list of
connection/stall markers (agent.py:237-246); an OpenRouter "is an embedding model" 400
matches none of them, so the real cause is discarded either way.
Persisted turns carry AgentTurn.stopped (agent.py:1178, "error" here), round-tripped to
the web as a bare string (services/web/src/lib/contracts.ts:455, stopped: z.string()), but
nothing in the web (stores/chat.ts, screens/ChatScreen.tsx) ever branches on stopped === "error" for a historical message — so even a turn that completes cleanly with stopped: "error" shows no visual distinction beyond the generic sentence baked into the reply
text.
The owner's screenshot shows neither the original text with an appended note nor an error
banner — just the bare partial with Copy/Regenerate. That's consistent with the SSE stream
dying before the note+done frames arrived: services/web/src/lib/sse.ts + stores/chat.ts:394-454
(consume) treat a broken mid-stream connection as "dropped", not "error"/"done"
(chat.ts:448,452), which hands off to reattachLoop (chat.ts:501+) — a re-attach that
exhausts its budget can give up without ever populating error for the user to see.
No ERROR-level log, no metric: agent.py:1161 and :1510 are both WARN, and nothing tracks a
failed/interrupted streaming turn as a counter. Core-app already defines Prometheus counters
elsewhere in this exact style (file_scan.py:37-46, e.g. epicurus_core_file_scan_fuse_trips_total), so the pattern to follow is established.
Fix
1 — model capability at the contract level (core, constraint #8)
add_saved_model (llm/routes.py:448-465) and set_default/set_embed_default
(:312-326) reject a mode mismatch with 400 before anything is persisted: no embedding model
as chat default, no chat model as embed default.
The web's chat ModelPicker (ChatScreen.tsx) and the "Hosted models" star-picker
(ModelsScreen.tsx:1734+) filter out embedding-only ids from the chat surfaces (and vice
versa in EmbedDefault, ModelsScreen.tsx:1458+), driven by the new mode/capability field
rather than a help-text warning.
LlmGateway.chat/stream/stream_chat (gateway.py:463-644) refuse a non-chat model before
calling litellm.acompletion, with a clear, model-actionable message — so a future picker gap
(or an operator hand-typing a hosted id) fails fast with a real reason instead of an opaque
provider 400 after retries.
2 — bring the facts recall store's self-heal up to the same reliability bar
Root-cause why UserFactStore._ensure (memory/facts.py:94-128) didn't reconcile on first
touch here — instrument or assert the isinstance(vectors_config, VectorParams) branch
(facts.py:117) so a response shape it doesn't recognize is a loud failure, not a silent current_dim = None no-op.
Never mark a collection _ensured (facts.py:128) unless the dimension was actually
confirmed to match afterward — a failed/skipped reconcile must stay un-cached so the next
call retries the fix instead of repeating the same doomed query forever.
Name the cause in the log: agent.py:1508-1515 should distinguish "embedding dimension
changed 768→4096; recall/save needs a rebuild" from a generic backend error, and surface that
state on the Models page (mirroring how the knowledge/notes dimension-change contract,
ADR-0132, reports itself) rather than only a WARN line.
Kubernetes parity: the reconcile is triggered lazily from request handling, not from
container startup, so both ADR-0134 runtime arms behave identically here — no change needed,
but call this out explicitly in the docs update below so it isn't assumed to be Docker-only.
3 — a failed streaming turn always tells the user something happened
In agent.py:1160-1186, when the "keep partial" branch runs, still emit a terminal error
event (in addition to — not instead of — the appended note and done), carrying the real banner text computed at agent.py:1162 (already scrubbed of connection-noise by _stream_failure_messages; drop only the OpenRouter user_id, which the operator doesn't
need). chat.ts's existing error handling (stores/chat.ts:416-419) then fires for this
case too instead of only the empty-partial one.
Web: render stopped === "error" distinctly on a persisted (historical) message, not only
on a live stream — an inline error affordance under the bubble with the reason and a
Regenerate action, so a reload or a re-attach that lands on "done" still shows what
happened.
Add a Prometheus counter for a failed/interrupted streaming turn (pattern: file_scan.py:37-46), e.g. epicurus_core_llm_stream_failures_total{tenant,reason}, so the
observability stack has something to alert on beyond a WARN line nobody is watching.
Tests
Part 1: a unit test asserting add_saved_model/set_default reject an embedding-mode model
id (and the embed-default routes reject a chat-mode one); a gateway test that chat/stream/stream_chat raise a clear, typed error for a non-chat model without ever
calling litellm.acompletion.
Part 2: a UserFactStore test against a fake/stub Qdrant client whose get_collection returns
a vectors-config shape the isinstance check doesn't recognize, asserting the store no longer
marks the collection _ensured and instead raises/logs loudly; a regression test that a real
dimension mismatch is fully healed by the second call at the latest.
Part 3: an agent-level test that a mid-stream exception with non-empty parts still yields a
terminal type="error" event before done; a web test (extending stores/chat.ts's existing
coverage) that a historical message with stopped: "error" renders the inline error
affordance.
Docs
docs/services/core-app.md — the model-roles section (near the saved-models/gateway routing
description) documents the new mode capability and the chat/embed rejection; the recall
section (:751-762, :1883-1897) documents the tightened _ensure invariant and what the
Models page now surfaces for a stuck dimension mismatch.
docs/reference/config.md — any new config key the fix introduces (none anticipated for parts
1–2; note if the new counter needs a scrape-interval-sensitive setting).
docs/reference/platform-api.md — the streaming contract: document the terminal error event
now being emitted alongside a retained partial turn, and the stopped field's meaning for
historical messages.
The Models settings page's own in-app copy (ModelsScreen.tsx) — replace the one-sided
"pick an embedding model there" warning with the real filtering described above.
Note: docs/reference/events.md's "Provider caveats" section (docs(reference): fill 'Provider caveats' and clear the last stale framing #852) is about module event
emitters (mail/calendar/tasks), not LLM providers — it is not the right home for an
"embedding models are not chat models" caveat; put that in docs/services/core-app.md instead.
What
Seen by the owner on the testing box (2026-09-13). As of
main@ d201052.Chat transcript (web shell): the owner writes "yeah, move them over there"; the assistant
replies with a single line — "Let me verify the current state first — the tree, any pending
suggestions, and the notes" — and the turn ends there. No further text, no tool calls, no
error banner. Just Copy / Regenerate under the partial reply.
Core-app log at the same moment (WARN, all
core-app):The owner had recently switched the tenant's embedding model to a hosted OpenRouter one
(
qwen/qwen3-embedding-8b, 4096 dims) through the gateway work of #868/#865. Three distinctdefects follow, filed together because they are one incident chain: a mis-scoped model picker
let an embedding model become the chat default, recall then silently lost memory instead of
healing, and the resulting mid-turn failure left the shell with a bare partial reply. Related:
#879 (search doesn't self-heal a dimension mismatch the way indexing does — the knowledge/notes
half of this same class of bug); #936 and #937 are the two other recent dogfood issues on this
milestone, filed separately (websearch degrade-detection and image-PDF OCR) — siblings, not
duplicates.
Where it lives
1 — the chat/completions call went to the embedding model
The saved-hosted-models list is capability-blind end to end, and nothing downstream ever checks
a model's chat-vs-embedding mode before handing it to
litellm.acompletion.services/web/src/screens/ChatScreen.tsx:699-716— the chatModelPicker's free-text hostedbox calls
chooseHosted(custom.trim())on submit, which both sets the model and persists it(
ChatScreen.tsx:642-645,save.mutate(id)→api.addSavedModel).services/core-app/src/epicurus_core_app/llm/routes.py:448-465(add_saved_model) validatesonly
is_hosted(model)(llm/providers.py:64-76— true for any<provider>/<name>string witha known alias); it never asks whether the id is chat-capable.
services/web/src/screens/ModelsScreen.tsx:1734-1820(SavedHostedModels) lists everysaved id under a star button wired to
setDefault.mutate→api.setGlobalDefault→PUT /platform/v1/llm/prefs/default.services/core-app/src/epicurus_core_app/llm/routes.py:312-318(set_default) writeswhatever string it's given straight into
llm_prefs.global_default— no validation at all.This rules out the "chat and embed prefs share a key" theory:
llm/prefs.py:35-37and:105-129keepglobal_defaultandembed_defaultas two distinct nullable columns; theynever collide.
services/core-app/src/epicurus_core_app/llm/gateway.py:197-203(effective_default) returnsthe stored value unchecked, and
chat()/stream()/stream_chat()(
gateway.py:463-521,522-644) hand the resolved candidate straight tolitellm.acompletion(...)with no pre-flight capability check — OpenRouter is the only thingin this call graph that ever rejects an embedding model for chat use.
gateway.py:977-1065(
show/_hosted_details) never reads LiteLLM's model-infomodefield (chatvsembeddingvs others in its cost map); it derives capabilities only from vision support andalways assumes
toolsfor a hosted id. For an id outside LiteLLM's static map — likely for anewer OpenRouter alias like
qwen/qwen3-embedding-8b—_note_unmapped(
gateway.py:909-922) logs once and the call defaults to["tools"]: capability-wiseindistinguishable from a genuine chat model.
SavedModel.capabilities(
llm/routes.py:136-137) round-trips exactly this value to the web, so the star-picker'sbadges can't tell the two apart either.
embedding side:
services/web/src/screens/ModelsScreen.tsx:1490-1493— "pick anembedding model there; a chat model will fail at embed time." Nothing warns the other
direction, and neither warning is backed by real capability data.
2 — recall silently degraded after the dimension change
agent/agent.py:1483-1515(_recall_within_budget) time-boxes the recall embed and swallowsany exception as
log.warning("recall skipped: backend error", ...)(line 1510) — the turnproceeds with no recalled facts, invisible to the owner.
This is not an uncontracted store, contrary to what the raw log suggests: the core's own
memory-fact store already carries a dimension-drift self-heal, older than #868 —
memory/facts.py:94-128(UserFactStore._ensure, #436/ADR-0074), documented atdocs/services/core-app.md:759-762and:1890-1897as reconciling "on first use each processlifetime". Both
save()(facts.py:224) andsearch()(facts.py:282, whichMemory.recall→UserFactStore.recall→searchall funnel through,memory/memory.py:110-112)call
_ensurebefore touching Qdrant.On paper the very first recall after the model switch should have reconciled the
<tenant>__factscollection (
scope_collection("facts", tenant),epicurus_core/tenancy.py:81-83) to 4096-d andevery later call should succeed. The log shows the opposite: the identical raw
UnexpectedResponse("expected dim: 768, got 4096") recurs three times over ~45s(12:37:27 / :38 / 12:38:12), each with a fast
elapsed_s(0.61 / 0.21 / 0.27) — far too quick tobe an in-flight re-embed of stored facts, and inconsistent with the reconcile ever actually
running. Two candidate mechanisms in
_ensureitself would produce exactly this symptom:facts.py:117—current_dim = vectors_config.size if isinstance(vectors_config, VectorParams) else None. If this ever evaluatesNone(aqdrant-clientresponse shape the check doesn'trecognize), the drift branch is skipped entirely — no reconcile happens.
facts.py:128—self._ensured.add(collection)runs unconditionally after theifblock,including the no-op path above. That permanently caches the collection as "reconciled" for the
rest of the process's life even though nothing was fixed, so every later
search()/save()repeats the identical fast 400 forever — matching the log pattern exactly.
Either way, the operator has no way to see "the recall store still needs a rebuild" versus "Qdrant
is unreachable" —
agent.py:1508-1515logs only the rawstr(exc)/error_typeat WARN, andnothing surfaces on the Models page (contrast the
knowledge/notesdimension-change contract,ADR-0132, which is drop-and-recreate rather than in-place re-embed and evidently holds up better
under a large dimension jump).
3 — a failed streaming turn leaves the UI with a partial reply and no error
agent/agent.py:1160-1186is the mid-stream exception handler wrapping the whole per-round toolloop (
agent.py:911-1159), so alitellm.BadRequestErrorraised by any round'sgateway.stream_chatcall (agent.py:927-928) lands here.not (partial.strip() or timeline),agent.py:1164), a real terminal event is emitted:yield AgentEvent(type="error", detail=banner)(agent.py:1167) — and the web already renders this correctly(
services/web/src/stores/chat.ts:416-419sets{error: detail, ...}).sentence before/around a tool call — the code takes the other branch (
agent.py:1169-1186).banner(the real reason: the raw litellm/OpenRouter message) is computed atagent.py:1162but is never used on this branch. Only a hardcoded generic
note—_STREAM_INTERRUPTED_MESSAGE= "The answer was interrupted before it finished. Please tryagain." (
agent.py:251) — gets appended as one moredelta(agent.py:1174), followed bydone(agent.py:1186). Noerror-typed event is ever emitted on this path, sochat.ts's working error banner is unreachable for exactly the failure mode this bug reports._stream_failure_messages(agent.py:254-265) only special-cases a short list ofconnection/stall markers (
agent.py:237-246); an OpenRouter "is an embedding model" 400matches none of them, so the real cause is discarded either way.
AgentTurn.stopped(agent.py:1178,"error"here), round-tripped tothe web as a bare string (
services/web/src/lib/contracts.ts:455,stopped: z.string()), butnothing in the web (
stores/chat.ts,screens/ChatScreen.tsx) ever branches onstopped === "error"for a historical message — so even a turn that completes cleanly withstopped: "error"shows no visual distinction beyond the generic sentence baked into the replytext.
banner — just the bare partial with Copy/Regenerate. That's consistent with the SSE stream
dying before the note+done frames arrived:
services/web/src/lib/sse.ts+stores/chat.ts:394-454(
consume) treat a broken mid-stream connection as"dropped", not"error"/"done"(
chat.ts:448,452), which hands off toreattachLoop(chat.ts:501+) — a re-attach thatexhausts its budget can give up without ever populating
errorfor the user to see.agent.py:1161and:1510are both WARN, and nothing tracks afailed/interrupted streaming turn as a counter. Core-app already defines Prometheus counters
elsewhere in this exact style (
file_scan.py:37-46, e.g.epicurus_core_file_scan_fuse_trips_total), so the pattern to follow is established.Fix
1 — model capability at the contract level (core, constraint #8)
mode(chat|embedding| …), sourced from LiteLLM'smodel-info map (
litellm.get_model_info(...)["mode"]) where mapped, plus the tenant'sexisting per-model override (
llm/saved_models.py'sSavedModelOverride, core-app,web: per-saved-model capability overrides — vision gating shouldn't trust litellm's static map #711) for theunmapped case — the same escape hatch already used for vision.
add_saved_model(llm/routes.py:448-465) andset_default/set_embed_default(
:312-326) reject a mode mismatch with 400 before anything is persisted: no embedding modelas chat default, no chat model as embed default.
ModelPicker(ChatScreen.tsx) and the "Hosted models" star-picker(
ModelsScreen.tsx:1734+) filter out embedding-only ids from the chat surfaces (and viceversa in
EmbedDefault,ModelsScreen.tsx:1458+), driven by the newmode/capability fieldrather than a help-text warning.
LlmGateway.chat/stream/stream_chat(gateway.py:463-644) refuse a non-chat model beforecalling
litellm.acompletion, with a clear, model-actionable message — so a future picker gap(or an operator hand-typing a hosted id) fails fast with a real reason instead of an opaque
provider 400 after retries.
2 — bring the facts recall store's self-heal up to the same reliability bar
UserFactStore._ensure(memory/facts.py:94-128) didn't reconcile on firsttouch here — instrument or assert the
isinstance(vectors_config, VectorParams)branch(
facts.py:117) so a response shape it doesn't recognize is a loud failure, not a silentcurrent_dim = Noneno-op._ensured(facts.py:128) unless the dimension was actuallyconfirmed to match afterward — a failed/skipped reconcile must stay un-cached so the next
call retries the fix instead of repeating the same doomed query forever.
agent.py:1508-1515should distinguish "embedding dimensionchanged 768→4096; recall/save needs a rebuild" from a generic backend error, and surface that
state on the Models page (mirroring how the knowledge/notes dimension-change contract,
ADR-0132, reports itself) rather than only a WARN line.
container startup, so both ADR-0134 runtime arms behave identically here — no change needed,
but call this out explicitly in the docs update below so it isn't assumed to be Docker-only.
3 — a failed streaming turn always tells the user something happened
agent.py:1160-1186, when the "keep partial" branch runs, still emit a terminalerrorevent (in addition to — not instead of — the appended note and
done), carrying the realbannertext computed atagent.py:1162(already scrubbed of connection-noise by_stream_failure_messages; drop only the OpenRouteruser_id, which the operator doesn'tneed).
chat.ts's existingerrorhandling (stores/chat.ts:416-419) then fires for thiscase too instead of only the empty-partial one.
stopped === "error"distinctly on a persisted (historical) message, not onlyon a live stream — an inline error affordance under the bubble with the reason and a
Regenerate action, so a reload or a re-attach that lands on
"done"still shows whathappened.
file_scan.py:37-46), e.g.epicurus_core_llm_stream_failures_total{tenant,reason}, so theobservability stack has something to alert on beyond a WARN line nobody is watching.
Tests
add_saved_model/set_defaultreject an embedding-mode modelid (and the embed-default routes reject a chat-mode one); a gateway test that
chat/stream/stream_chatraise a clear, typed error for a non-chat model without evercalling
litellm.acompletion.UserFactStoretest against a fake/stub Qdrant client whoseget_collectionreturnsa vectors-config shape the
isinstancecheck doesn't recognize, asserting the store no longermarks the collection
_ensuredand instead raises/logs loudly; a regression test that a realdimension mismatch is fully healed by the second call at the latest.
partsstill yields aterminal
type="error"event beforedone; a web test (extendingstores/chat.ts's existingcoverage) that a historical message with
stopped: "error"renders the inline erroraffordance.
Docs
docs/services/core-app.md— the model-roles section (near the saved-models/gateway routingdescription) documents the new
modecapability and the chat/embed rejection; the recallsection (
:751-762,:1883-1897) documents the tightened_ensureinvariant and what theModels page now surfaces for a stuck dimension mismatch.
docs/reference/config.md— any new config key the fix introduces (none anticipated for parts1–2; note if the new counter needs a scrape-interval-sensitive setting).
docs/reference/platform-api.md— the streaming contract: document the terminalerroreventnow being emitted alongside a retained partial turn, and the
stoppedfield's meaning forhistorical messages.
ModelsScreen.tsx) — replace the one-sided"pick an embedding model there" warning with the real filtering described above.
docs/reference/events.md's "Provider caveats" section (docs(reference): fill 'Provider caveats' and clear the last stale framing #852) is about module eventemitters (mail/calendar/tasks), not LLM providers — it is not the right home for an
"embedding models are not chat models" caveat; put that in
docs/services/core-app.mdinstead.core-appMINOR,webMINOR.