Skip to content

[Bug]: Tool calls are correctly received but never executed for custom OpenAI-compatible endpoint (Bifrost gateway → vLLM) #14435

Description

@Feuerhamster

Summary

When using a custom OpenAI-compatible endpoint (a self-hosted Bifrost gateway in front of vLLM, serving GLM-5.2 and gpt-oss-120b), LibreChat correctly sends tools in the request and the upstream model correctly responds with a valid, complete tool_calls payload (verified via a logging man-in-the-middle proxy). However, LibreChat never executes the tool call and never makes the expected second LLM call with the tool result — it simply ends the turn immediately, resulting in an empty/near-empty final assistant message (completionTokens: 4, no visible content).

This reproduces identically for:

  • web_search (built-in tool)
  • an MCP server's tools (memos)
  • multiple models behind the same custom endpoint (GLM-5.2, gpt-oss-120b, Gemma)
  • both stream: true and non-streaming title-generation calls
  • LibreChat v0.8.7 (latest) and v0.8.6 (downgrade tested, same result)

Environment

  • LibreChat version: v0.8.7 (also reproduced on v0.8.6)
  • Deployment: Docker Compose, self-hosted
  • Endpoint type: endpoints.custom (OpenAI-compatible), pointing at a self-hosted Bifrost gateway (multi-provider LLM proxy), which forwards to a vLLM cluster
  • Models tested: vllm/release/glm-5-2, vllm/release/gpt-oss-120b, vllm/release/gemma-4-31b-it
  • Web search config: searchProvider: searxng (self-hosted SearXNG) and searchProvider: tavily (both tested, same result) + self-hosted Firecrawl-compatible scraper
  • MCP: one streamable-http MCP server (memos) also exhibits the same "tool never executes" symptom

Steps to Reproduce

  1. Configure a custom endpoint per the docs, pointing at any OpenAI-compatible backend that supports tool calling (verified working directly via curl, see evidence below).
  2. Enable the built-in Web Search feature (webSearch config, any provider).
  3. Send a message that requires current information, e.g. "What were yesterday's Formula 1 results?"
  4. Observe: the assistant message ends up empty or with a couple of throwaway tokens. No search is performed. No error is shown to the user.

Expected Behavior

  1. LLM responds with tool_calls (confirmed happening correctly upstream).
  2. LibreChat executes the corresponding tool (web_search, or the MCP tool).
  3. LibreChat sends a second request to the LLM with the tool result appended as a role: "tool" message.
  4. The LLM produces the final, content-bearing answer.

Actual Behavior

Only step 1 happens. LibreChat's own debug logs show a single [agents:graph] Invoking LLM → LLM call complete → Emitting FINAL event sequence per turn, with completionTokens: 4 and no second LLM invocation and no visible tool execution logs. The chat then simply shows an empty response.

Evidence

To rule out any misconfiguration or upstream (gateway/model) fault, we inserted a transparent logging reverse proxy between LibreChat and the Bifrost gateway (baseURL pointed at the proxy, which forwarded 1:1 to the real gateway and logged both request and response bodies, including full SSE stream reconstruction).

Request received from LibreChat (tools array present, well-formed, single web_search function):

{
  "model": "vllm/release/glm-5-2",
  "stream": true,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "web_search",
        "description": "...",
        "parameters": {
          "type": "object",
          "properties": {
            "query": { "type": "string", "description": "..." },
            "date": { "type": "string", "enum": ["h","d","w","m","y"] },
            "country": { "type": "string", "description": "..." },
            "images": { "type": "boolean" },
            "videos": { "type": "boolean" },
            "news": { "type": "boolean" }
          },
          "required": ["query"]
        }
      }
    }
  ],
  "messages": [ /* system + user message */ ]
}

Response received from the upstream model (reconstructed from the SSE stream, reasoning deltas omitted for brevity): a fully-formed tool call and a correct finish reason:

{
  "choices": [
    {
      "index": 0,
      "finish_reason": "tool_calls",
      "delta": {}
    }
  ],
  "usage": { "prompt_tokens": 1287, "completion_tokens": 99, "total_tokens": 1386 }
}

The accumulated tool_calls delta stream assembles into a fully valid call:

{
  "id": "call_...",
  "type": "function",
  "function": {
    "name": "web_search",
    "arguments": "{\"query\": \"Formel 1 Ergebnisse 24. Juli 2026\", \"date\": \"d\", \"country\": \"de\"}"
  }
}

This was independently reproduced multiple times, for both GLM-5.2 and gpt-oss-120b, with and without a "reasoning"/"reasoning_details" field present in intermediate chunks (we tested stripping reasoning/reasoning_details from every chunk via the proxy to rule out a reasoning-field-name mismatch — no change in behavior).

Meanwhile, LibreChat's own debug log for the same turn shows no second LLM call and no tool-execution log line at all:

debug: [agents:graph] Invoking LLM { messageCount: 1, provider: "openAI", ... }
debug: [agents:graph] LLM call complete (1.01s) { messageCount: 1, ... }
debug: [spendTokens] ... completionTokens: 4
debug: [ResumableAgentController] Emitting FINAL event { wasAbortedBeforeComplete: false, ... }

No error is logged. wasAbortedBeforeComplete is false. The turn is simply treated as complete after the first (tool-call) response, without executing the tool or looping back to the model.

Additional data point (MCP)

The same pattern occurs with an MCP server (memos, streamable-http, 19 tools listed successfully via /api/mcp/tools). Asking the agent to use an MCP tool produces a short assistant message announcing intent ("Ich rufe deine aktuellen Notizen ab...") but the tool is never actually invoked, and the log shows:

debug: [getMCPTools] No tools found for server memos

immediately before the (tool-less) LLM call — even though the same server correctly lists 19 tools via the REST API moments earlier in the same session.

What we ruled out

Suspected Area

Given the evidence, the issue appears to be in the agent orchestration logic that decides whether to execute a tool call and loop back to the LLM for a custom (endpointType: "custom") OpenAI-compatible endpoint specifically — possibly related to how the tool_calls finish reason or accumulated tool call arguments are detected for this endpoint type in @librechat/agents, as opposed to native openAI/azureOpenAI endpoint types (untested by us, we don't have those configured).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions