Repository navigation
fix: use the Chat Completions adapter so streamed tool loops work on OpenAI-compatible gateways - #93
matthewhand wants to merge 1 commit into
Conversation
The realtime path was permanently broken on Node: the local WS proxy used Cloudflare Workers APIs (WebSocketPair, response.webSocket) inside @hono/node-server, so every browser upgrade 502'd and connectAgent hung on "Loading conversation…". Drop the proxy and let browsers join Intelligence's managed realtime WS directly, as OpenBot already does. While verifying end to end, the LiteLLM gateway rejected follow-up turns with 400 "Thinking level MINIMAL is not supported" (its Responses→Vertex bridge injects a default thinking level) and its provider chain choked on strict no-arg tool schemas and replayed tool-call ids. Switch the engine to the Chat Completions adapter and send an explicit reasoning effort (OPENAI_REASONING_EFFORT, default low) so the gateway never injects one. The model fixture now serves the Chat Completions protocol to match. Also: single-source the Intelligence user id as config.intelligenceUserId (INTELLIGENCE_USER_ID, default local-user) used by both identifyUser and the main-thread endpoint; honor skipAccessKey in live-mode session auth; and fix a missing import in BrowserConsole.web.tsx. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
|
The Chat Completions adapter correction looks useful, but head If a separate local-development option is needed, isolate it in its own change with enforced loopback-only operation, explicit documentation, and tests that non-loopback live startup and missing/invalid session keys are rejected. Please add a regression test retaining the normal live session rejection before merging this PR. Validation was static review of the current diff and session implementation; runtime tests were not rerun. Changed code: apps/server/src/auth.ts:18. |
jerelvelarde
left a comment
There was a problem hiding this comment.
The gateway benefit overlaps merged #80's opt-in Chat Completions adapter. This change instead globally replaces the default Responses path and adds reasoning/identity changes; isolate and document any remaining incremental value while preserving the opt-in behavior. Five configuration and four model-worker tests passed, but the reproduced live-authentication bypass below blocks approval. Integrate the engine/tanstack-agent.ts production conflict against current main.
| async session(accessKey?: string) { | ||
| if ( | ||
| this.config.mode === "live" && | ||
| !this.config.skipAccessKey && |
There was a problem hiding this comment.
[P1] Remove this unrelated live authentication bypass. OPENMUSE_SKIP_ACCESS_KEY=true disables both the live access-key check and startup validation, without a replacement authentication boundary. Reproduced with the real application: unauthenticated POST /api/session with {} returned 200, and its issued token accessed /api/workspace with 200; disabling the flag returned 401 for keyless authentication. Live deployments can bind publicly, and the gateway PR does not document this access change. Remove the bypass from this fix and preserve live access-key enforcement.
Problem
On OpenAI-compatible gateways (LiteLLM in front of Vertex/Groq), the Responses-API adapter breaks on every multi-turn tool loop:
thinking_level: MINIMALon follow-up turns; Vertex-backed models reject it with400 "Thinking level MINIMAL is not supported".chatcmpl-tool-…ids back through the Responses bridge, which rejects them:Invalid 'input[1].id': … Expected an ID that begins with 'fc'.requiredwithoutproperties).Turn 1 streams fine; the tool-result follow-up turn dies. Same failure shape regardless of model.
Fix
Switch
adapter()to the Chat Completions adapter (openaiChatCompletions) — the wire format these gateways handle natively — and send an explicitreasoning_effort(newOPENAI_REASONING_EFFORTenv, defaultlow) so the gateway never injects its own default. Anthropic/Gemini paths unchanged.tests/helpers/model.tsserved the Responses SSE protocol; it now serves Chat Completions (including the in-stream{error}chunk the OpenAI SDK throws on, preserving the not-retried semantics the tests assert).Testing
main;tsc --noEmitclean.Adapted from a deployment fix; the same conclusion was reached independently in sunshaoan0808/openmuse ("Force Chat Completions so tool calls work on vLLM/LiteLLM").
🤖 Generated with Codebuff