Skip to content

A streaming turn dies silently after the embedding model changes #944

Description

@baakhoff

What

Seen by the owner on the testing box (2026-09-13). As of main @ d201052.

Chat transcript (web shell): the owner writes "yeah, move them over there"; the assistant
replies with a single line — "Let me verify the current state first — the tree, any pending
suggestions, and the notes" — and the turn ends there. No further text, no tool calls, no
error banner. Just Copy / Regenerate under the partial reply.

Core-app log at the same moment (WARN, all core-app):

12:37:27.026 WARN recall skipped: backend error
  {"error": "Unexpected Response: 400 (Bad Request)\n...Vector dimension error: expected dim: 768, got 4096",
   "error_type": "UnexpectedResponse", "elapsed_s": 0.61}
12:37:27.252 WARN streaming turn failed
  {"error": "litellm.BadRequestError: OpenrouterException - {\"error\":{\"message\":\"qwen/qwen3-embedding-8b
   is an embedding model and cannot be used with the chat/completions endpoint. Use the /embeddings endpoint
   instead.\",\"code\":400}}...\" LiteLLM Retried: 2 times"}
12:37:38.512 WARN recall skipped: backend error  (same 768-vs-4096 payload, elapsed_s 0.21)
12:38:12.441 WARN recall skipped: backend error  (same, elapsed_s 0.27)

The owner had recently switched the tenant's embedding model to a hosted OpenRouter one
(qwen/qwen3-embedding-8b, 4096 dims) through the gateway work of #868/#865. Three distinct
defects follow, filed together because they are one incident chain: a mis-scoped model picker
let an embedding model become the chat default, recall then silently lost memory instead of
healing, and the resulting mid-turn failure left the shell with a bare partial reply. Related:
#879 (search doesn't self-heal a dimension mismatch the way indexing does — the knowledge/notes
half of this same class of bug); #936 and #937 are the two other recent dogfood issues on this
milestone, filed separately (websearch degrade-detection and image-PDF OCR) — siblings, not
duplicates.

Where it lives

1 — the chat/completions call went to the embedding model

The saved-hosted-models list is capability-blind end to end, and nothing downstream ever checks
a model's chat-vs-embedding mode before handing it to litellm.acompletion.

  • services/web/src/screens/ChatScreen.tsx:699-716 — the chat ModelPicker's free-text hosted
    box calls chooseHosted(custom.trim()) on submit, which both sets the model and persists it
    (ChatScreen.tsx:642-645, save.mutate(id) → api.addSavedModel).
  • services/core-app/src/epicurus_core_app/llm/routes.py:448-465 (add_saved_model) validates
    only is_hosted(model) (llm/providers.py:64-76 — true for any <provider>/<name> string with
    a known alias); it never asks whether the id is chat-capable.
  • services/web/src/screens/ModelsScreen.tsx:1734-1820 (SavedHostedModels) lists every
    saved id under a star button wired to setDefault.mutate → api.setGlobalDefault →
    PUT /platform/v1/llm/prefs/default.
  • services/core-app/src/epicurus_core_app/llm/routes.py:312-318 (set_default) writes
    whatever string it's given straight into llm_prefs.global_default — no validation at all.
    This rules out the "chat and embed prefs share a key" theory: llm/prefs.py:35-37 and
    :105-129 keep global_default and embed_default as two distinct nullable columns; they
    never collide.
  • services/core-app/src/epicurus_core_app/llm/gateway.py:197-203 (effective_default) returns
    the stored value unchecked, and chat()/stream()/stream_chat()
    (gateway.py:463-521, 522-644) hand the resolved candidate straight to
    litellm.acompletion(...) with no pre-flight capability check — OpenRouter is the only thing
    in this call graph that ever rejects an embedding model for chat use.
  • The capability data needed to prevent this doesn't exist: gateway.py:977-1065
    (show/_hosted_details) never reads LiteLLM's model-info mode field (chat vs
    embedding vs others in its cost map); it derives capabilities only from vision support and
    always assumes tools for a hosted id. For an id outside LiteLLM's static map — likely for a
    newer OpenRouter alias like qwen/qwen3-embedding-8b — _note_unmapped
    (gateway.py:909-922) logs once and the call defaults to ["tools"]: capability-wise
    indistinguishable from a genuine chat model. SavedModel.capabilities
    (llm/routes.py:136-137) round-trips exactly this value to the web, so the star-picker's
    badges can't tell the two apart either.
  • The only place this is even documented as a hazard is a sentence of help text on the
    embedding side: services/web/src/screens/ModelsScreen.tsx:1490-1493 — "pick an
    embedding model there; a chat model will fail at embed time." Nothing warns the other
    direction, and neither warning is backed by real capability data.

2 — recall silently degraded after the dimension change

agent/agent.py:1483-1515 (_recall_within_budget) time-boxes the recall embed and swallows
any exception as log.warning("recall skipped: backend error", ...) (line 1510) — the turn
proceeds with no recalled facts, invisible to the owner.

This is not an uncontracted store, contrary to what the raw log suggests: the core's own
memory-fact store already carries a dimension-drift self-heal, older than #868 —
memory/facts.py:94-128 (UserFactStore._ensure, #436/ADR-0074), documented at
docs/services/core-app.md:759-762 and :1890-1897 as reconciling "on first use each process
lifetime". Both save() (facts.py:224) and search() (facts.py:282, which
Memory.recall → UserFactStore.recall → search all funnel through, memory/memory.py:110-112)
call _ensure before touching Qdrant.

On paper the very first recall after the model switch should have reconciled the <tenant>__facts
collection (scope_collection("facts", tenant), epicurus_core/tenancy.py:81-83) to 4096-d and
every later call should succeed. The log shows the opposite: the identical raw
UnexpectedResponse ("expected dim: 768, got 4096") recurs three times over ~45s
(12:37:27 / :38 / 12:38:12), each with a fast elapsed_s (0.61 / 0.21 / 0.27) — far too quick to
be an in-flight re-embed of stored facts, and inconsistent with the reconcile ever actually
running. Two candidate mechanisms in _ensure itself would produce exactly this symptom:

  • facts.py:117 — current_dim = vectors_config.size if isinstance(vectors_config, VectorParams) else None. If this ever evaluates None (a qdrant-client response shape the check doesn't
    recognize), the drift branch is skipped entirely — no reconcile happens.
  • facts.py:128 — self._ensured.add(collection) runs unconditionally after the if block,
    including the no-op path above. That permanently caches the collection as "reconciled" for the
    rest of the process's life even though nothing was fixed, so every later search()/save()
    repeats the identical fast 400 forever — matching the log pattern exactly.

Either way, the operator has no way to see "the recall store still needs a rebuild" versus "Qdrant
is unreachable" — agent.py:1508-1515 logs only the raw str(exc) / error_type at WARN, and
nothing surfaces on the Models page (contrast the knowledge/notes dimension-change contract,
ADR-0132, which is drop-and-recreate rather than in-place re-embed and evidently holds up better
under a large dimension jump).

3 — a failed streaming turn leaves the UI with a partial reply and no error

agent/agent.py:1160-1186 is the mid-stream exception handler wrapping the whole per-round tool
loop (agent.py:911-1159), so a litellm.BadRequestError raised by any round's
gateway.stream_chat call (agent.py:927-928) lands here.

  • When nothing was produced before the failure (not (partial.strip() or timeline),
    agent.py:1164), a real terminal event is emitted: yield AgentEvent(type="error", detail=banner) (agent.py:1167) — and the web already renders this correctly
    (services/web/src/stores/chat.ts:416-419 sets {error: detail, ...}).
  • When there is partial content or tool activity — the owner's exact scenario, a lead-in
    sentence before/around a tool call — the code takes the other branch (agent.py:1169-1186).
    banner (the real reason: the raw litellm/OpenRouter message) is computed at agent.py:1162
    but is never used on this branch. Only a hardcoded generic note —
    _STREAM_INTERRUPTED_MESSAGE = "The answer was interrupted before it finished. Please try
    again." (agent.py:251) — gets appended as one more delta (agent.py:1174), followed by
    done (agent.py:1186). No error-typed event is ever emitted on this path, so
    chat.ts's working error banner is unreachable for exactly the failure mode this bug reports.
    _stream_failure_messages (agent.py:254-265) only special-cases a short list of
    connection/stall markers (agent.py:237-246); an OpenRouter "is an embedding model" 400
    matches none of them, so the real cause is discarded either way.
  • Persisted turns carry AgentTurn.stopped (agent.py:1178, "error" here), round-tripped to
    the web as a bare string (services/web/src/lib/contracts.ts:455, stopped: z.string()), but
    nothing in the web (stores/chat.ts, screens/ChatScreen.tsx) ever branches on
    stopped === "error" for a historical message — so even a turn that completes cleanly with
    stopped: "error" shows no visual distinction beyond the generic sentence baked into the reply
    text.
  • The owner's screenshot shows neither the original text with an appended note nor an error
    banner — just the bare partial with Copy/Regenerate. That's consistent with the SSE stream
    dying before the note+done frames arrived: services/web/src/lib/sse.ts + stores/chat.ts:394-454
    (consume) treat a broken mid-stream connection as "dropped", not "error"/"done"
    (chat.ts:448,452), which hands off to reattachLoop (chat.ts:501+) — a re-attach that
    exhausts its budget can give up without ever populating error for the user to see.
  • No ERROR-level log, no metric: agent.py:1161 and :1510 are both WARN, and nothing tracks a
    failed/interrupted streaming turn as a counter. Core-app already defines Prometheus counters
    elsewhere in this exact style (file_scan.py:37-46, e.g.
    epicurus_core_file_scan_fuse_trips_total), so the pattern to follow is established.

Fix

1 — model capability at the contract level (core, constraint #8)

  1. Give a saved/known model a real mode (chat | embedding | …), sourced from LiteLLM's
    model-info map (litellm.get_model_info(...)["mode"]) where mapped, plus the tenant's
    existing per-model override (llm/saved_models.py's SavedModelOverride, core-app,web: per-saved-model capability overrides — vision gating shouldn't trust litellm's static map #711) for the
    unmapped case — the same escape hatch already used for vision.
  2. add_saved_model (llm/routes.py:448-465) and set_default/set_embed_default
    (:312-326) reject a mode mismatch with 400 before anything is persisted: no embedding model
    as chat default, no chat model as embed default.
  3. The web's chat ModelPicker (ChatScreen.tsx) and the "Hosted models" star-picker
    (ModelsScreen.tsx:1734+) filter out embedding-only ids from the chat surfaces (and vice
    versa in EmbedDefault, ModelsScreen.tsx:1458+), driven by the new mode/capability field
    rather than a help-text warning.
  4. LlmGateway.chat/stream/stream_chat (gateway.py:463-644) refuse a non-chat model before
    calling litellm.acompletion, with a clear, model-actionable message — so a future picker gap
    (or an operator hand-typing a hosted id) fails fast with a real reason instead of an opaque
    provider 400 after retries.

2 — bring the facts recall store's self-heal up to the same reliability bar

  1. Root-cause why UserFactStore._ensure (memory/facts.py:94-128) didn't reconcile on first
    touch here — instrument or assert the isinstance(vectors_config, VectorParams) branch
    (facts.py:117) so a response shape it doesn't recognize is a loud failure, not a silent
    current_dim = None no-op.
  2. Never mark a collection _ensured (facts.py:128) unless the dimension was actually
    confirmed to match afterward — a failed/skipped reconcile must stay un-cached so the next
    call retries the fix instead of repeating the same doomed query forever.
  3. Name the cause in the log: agent.py:1508-1515 should distinguish "embedding dimension
    changed 768→4096; recall/save needs a rebuild" from a generic backend error, and surface that
    state on the Models page (mirroring how the knowledge/notes dimension-change contract,
    ADR-0132, reports itself) rather than only a WARN line.
  4. Kubernetes parity: the reconcile is triggered lazily from request handling, not from
    container startup, so both ADR-0134 runtime arms behave identically here — no change needed,
    but call this out explicitly in the docs update below so it isn't assumed to be Docker-only.

3 — a failed streaming turn always tells the user something happened

  1. In agent.py:1160-1186, when the "keep partial" branch runs, still emit a terminal error
    event (in addition to — not instead of — the appended note and done), carrying the real
    banner text computed at agent.py:1162 (already scrubbed of connection-noise by
    _stream_failure_messages; drop only the OpenRouter user_id, which the operator doesn't
    need). chat.ts's existing error handling (stores/chat.ts:416-419) then fires for this
    case too instead of only the empty-partial one.
  2. Web: render stopped === "error" distinctly on a persisted (historical) message, not only
    on a live stream — an inline error affordance under the bubble with the reason and a
    Regenerate action, so a reload or a re-attach that lands on "done" still shows what
    happened.
  3. Add a Prometheus counter for a failed/interrupted streaming turn (pattern:
    file_scan.py:37-46), e.g. epicurus_core_llm_stream_failures_total{tenant,reason}, so the
    observability stack has something to alert on beyond a WARN line nobody is watching.

Tests

  • Part 1: a unit test asserting add_saved_model/set_default reject an embedding-mode model
    id (and the embed-default routes reject a chat-mode one); a gateway test that
    chat/stream/stream_chat raise a clear, typed error for a non-chat model without ever
    calling litellm.acompletion.
  • Part 2: a UserFactStore test against a fake/stub Qdrant client whose get_collection returns
    a vectors-config shape the isinstance check doesn't recognize, asserting the store no longer
    marks the collection _ensured and instead raises/logs loudly; a regression test that a real
    dimension mismatch is fully healed by the second call at the latest.
  • Part 3: an agent-level test that a mid-stream exception with non-empty parts still yields a
    terminal type="error" event before done; a web test (extending stores/chat.ts's existing
    coverage) that a historical message with stopped: "error" renders the inline error
    affordance.

Docs

  • docs/services/core-app.md — the model-roles section (near the saved-models/gateway routing
    description) documents the new mode capability and the chat/embed rejection; the recall
    section (:751-762, :1883-1897) documents the tightened _ensure invariant and what the
    Models page now surfaces for a stuck dimension mismatch.
  • docs/reference/config.md — any new config key the fix introduces (none anticipated for parts
    1–2; note if the new counter needs a scrape-interval-sensitive setting).
  • docs/reference/platform-api.md — the streaming contract: document the terminal error event
    now being emitted alongside a retained partial turn, and the stopped field's meaning for
    historical messages.
  • The Models settings page's own in-app copy (ModelsScreen.tsx) — replace the one-sided
    "pick an embedding model there" warning with the real filtering described above.
  • Note: docs/reference/events.md's "Provider caveats" section (docs(reference): fill 'Provider caveats' and clear the last stale framing #852) is about module event
    emitters (mail/calendar/tasks), not LLM providers — it is not the right home for an
    "embedding models are not chat models" caveat; put that in docs/services/core-app.md instead.

core-app MINOR, web MINOR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions