Tool calls are correctly received but never executed for custom OpenAI-compatible endpoint (Bifrost gateway → vLLM) #14822
Replies: 3 comments 1 reply
|
We reproduced a closely related failure on LibreChat v0.8.7 with a saved Agent using a custom OpenAI-compatible endpoint:
Our temporary workaround is to use LibreChat native provider Could LibreChat either fix the custom endpoint Agent tool loop, or make the native OpenAI transport/reverse-proxy route configurable per endpoint/agent? This would let self-hosters retain OpenAI-compatible gateways without patching code or migrating agents between provider types. |
|
Disclosure: TrustGate / NeuralTrust DevRel. This pattern (OpenAI-compatible gateway → vLLM) often breaks tool calls in one of three places:
If Bifrost logs show a full assistant message with We run an OSS Agent Gateway in front of OpenAI-compatible backends + MCP (Go, Apache-2.0): https://github.com/NeuralTrust/TrustGate — happy to compare notes on how LibreChat expects tool payloads shaped if useful. |
|
Same symptom here on v0.8.8-rc3 with LocalAI (llama-cpp backend, gemma-4-31b-it and other models), so this isn't specific to Bifrost/vLLM. I traced it to the root cause and have a one-line fix confirmed working. Root causeBackends that stream reasoning emit delta.reasoning chunks before the first chunk that carries role: "assistant". @langchain/openai (1.5.8, pinned by @librechat/agents 3.9.0) converts each delta in dist/converters/completions.js:
With Repro (LocalAI)Same request, stream: true, with tools → first chunk: Fix (confirmed)Defaulting the role to assistant — which is what the Python LangChain client already does — fixes it, with streaming, web search and file search working again:
(patched in node_modules/@langchain/openai/dist/converters/completions.{js,cjs}). Arguably the backends should send role on the first chunk too, but every OpenAI-compatible server that separates reasoning seems to get this wrong, so defaulting in the converter is the robust fix. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
When using a custom OpenAI-compatible endpoint (a self-hosted Bifrost gateway in front of vLLM, serving GLM-5.2 and gpt-oss-120b), LibreChat correctly sends
toolsin the request and the upstream model correctly responds with a valid, completetool_callspayload (verified via a logging man-in-the-middle proxy). However, LibreChat never executes the tool call and never makes the expected second LLM call with the tool result — it simply ends the turn immediately, resulting in an empty/near-empty final assistant message (completionTokens: 4, no visible content).This reproduces identically for:
web_search(built-in tool)memos)stream: trueand non-streaming title-generation callsv0.8.7(latest) andv0.8.6(downgrade tested, same result)Environment
endpoints.custom(OpenAI-compatible), pointing at a self-hosted Bifrost gateway (multi-provider LLM proxy), which forwards to a vLLM clustervllm/release/glm-5-2,vllm/release/gpt-oss-120b,vllm/release/gemma-4-31b-itsearchProvider: searxng(self-hosted SearXNG) andsearchProvider: tavily(both tested, same result) + self-hosted Firecrawl-compatible scrapermemos) also exhibits the same "tool never executes" symptomSteps to Reproduce
curl, see evidence below).webSearchconfig, any provider).Expected Behavior
tool_calls(confirmed happening correctly upstream).web_search, or the MCP tool).role: "tool"message.Actual Behavior
Only step 1 happens. LibreChat's own debug logs show a single
[agents:graph] Invoking LLM→LLM call complete→Emitting FINAL eventsequence per turn, withcompletionTokens: 4and no second LLM invocation and no visible tool execution logs. The chat then simply shows an empty response.Evidence
To rule out any misconfiguration or upstream (gateway/model) fault, we inserted a transparent logging reverse proxy between LibreChat and the Bifrost gateway (
baseURLpointed at the proxy, which forwarded 1:1 to the real gateway and logged both request and response bodies, including full SSE stream reconstruction).Request received from LibreChat (
toolsarray present, well-formed, singleweb_searchfunction):{ "model": "vllm/release/glm-5-2", "stream": true, "tools": [ { "type": "function", "function": { "name": "web_search", "description": "...", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "..." }, "date": { "type": "string", "enum": ["h","d","w","m","y"] }, "country": { "type": "string", "description": "..." }, "images": { "type": "boolean" }, "videos": { "type": "boolean" }, "news": { "type": "boolean" } }, "required": ["query"] } } } ], "messages": [ /* system + user message */ ] }Response received from the upstream model (reconstructed from the SSE stream, reasoning deltas omitted for brevity): a fully-formed tool call and a correct finish reason:
{ "choices": [ { "index": 0, "finish_reason": "tool_calls", "delta": {} } ], "usage": { "prompt_tokens": 1287, "completion_tokens": 99, "total_tokens": 1386 } }The accumulated
tool_callsdelta stream assembles into a fully valid call:{ "id": "call_...", "type": "function", "function": { "name": "web_search", "arguments": "{\"query\": \"Formel 1 Ergebnisse 24. Juli 2026\", \"date\": \"d\", \"country\": \"de\"}" } }This was independently reproduced multiple times, for both GLM-5.2 and gpt-oss-120b, with and without a "reasoning"/"reasoning_details" field present in intermediate chunks (we tested stripping
reasoning/reasoning_detailsfrom every chunk via the proxy to rule out a reasoning-field-name mismatch — no change in behavior).Meanwhile, LibreChat's own debug log for the same turn shows no second LLM call and no tool-execution log line at all:
No error is logged.
wasAbortedBeforeCompleteisfalse. The turn is simply treated as complete after the first (tool-call) response, without executing the tool or looping back to the model.Additional data point (MCP)
The same pattern occurs with an MCP server (
memos, streamable-http, 19 tools listed successfully via/api/mcp/tools). Asking the agent to use an MCP tool produces a short assistant message announcing intent ("Ich rufe deine aktuellen Notizen ab...") but the tool is never actually invoked, and the log shows:immediately before the (tool-less) LLM call — even though the same server correctly lists 19 tools via the REST API moments earlier in the same session.
What we ruled out
curl(non-streaming and streaming) against the same Bifrost endpoint/model — both GLM-5.2 and gpt-oss-120b return well-formedtool_callswithfinish_reason: "tool_calls"immediately, no retries needed.searchProvider: searxngandsearchProvider: tavily; reproduced identically with different scraper backends.reasoningvsreasoning_contentfield naming (cf. [Bug]: Reasoning content not rendered for custom OpenAI-compatible endpoints (vLLM, SGLang) — hardcoded to deprecated reasoning_content field #12775, [Bug]: v0.8.7-rc1 stops rendering reasoning and breaks tool calling for a custom OpenAI-compatible endpoint (works on v0.8.6, backend verified correct) #13844): testedcustomParams.reasoningKey, and separately tested a proxy that stripsreasoning/reasoning_detailsentirely from every streamed chunk before it reaches LibreChat — no change in behavior.v0.8.7andv0.8.6.Suspected Area
Given the evidence, the issue appears to be in the agent orchestration logic that decides whether to execute a tool call and loop back to the LLM for a custom (
endpointType: "custom") OpenAI-compatible endpoint specifically — possibly related to how thetool_callsfinish reason or accumulated tool call arguments are detected for this endpoint type in@librechat/agents, as opposed to nativeopenAI/azureOpenAIendpoint types (untested by us, we don't have those configured).All reactions