Skip to content

[Bug / Potential Fix] DeepSeek V4 Flash/Pro: reasoning_content must be passed back to the API (400) on assistant messages with tool_calls #878

Description

@nickmesen

[Bug]: DeepSeek V4 — 400 error caused by reasoning_content missing on edge-case assistant messages

Author's note: This report builds on the foundational DeepSeek V4 support added by Jatmn in commit ff2a380 (Add DeepSeek V4 flash/pro support and DeepSeek thinking compatibility #877). That commit appears to have provided the core infrastructure: DeepSeek URL detection, thinking: { type: 'enabled' } support, and the preserveReasoningContent mechanism. We are very grateful for that work.

What we found: After testing extensively with deepseek-v4-pro, we found what appear to be four edge cases not covered by ff2a380 where reasoning_content was not being set, which in our DeepSeek reproductions led to 400 errors in multi-turn conversations with tool use. The original commit seems to handle the happy path well (assistant messages that do contain thinking blocks), but may miss assistant messages without thinking blocks, string-content messages, and synthetic interrupt messages.

Our validation for this report was performed against DeepSeek. While the proposed explanation and impact assessment seem reasonable, we are not fully certain there are no secondary effects in other paths or providers, so review from maintainers with deeper context would be greatly appreciated.


Summary

When using OpenClaude with DeepSeek V4 (deepseek-v4-flash, deepseek-v4-pro) as an OpenAI-compatible provider, conversations that involve tool calls can fail with a 400 invalid_request_error. The error message is:

The `reasoning_content` in the thinking mode must be passed back to the API.

This appears to happen because:

  1. DeepSeek V4 enables thinking mode by default, which emits reasoning_content on assistant messages.
  2. Commit ff2a380 added the infrastructure to capture thinking blocks and re-attach them as reasoning_content.
  3. However, what appear to be four edge cases were not covered, causing reasoning_content to be undefined on some assistant messages.
  4. In the DeepSeek scenarios we validated, the request is rejected if an assistant message in thinking mode omits reasoning_content.

This issue appears to affect tool-calling conversations with DeepSeek V4, especially subagents / Task tool invocations where tools are used repeatedly across turns.


Environment

Component Version / Setting
OpenClaude v0.6.0
Provider openai-compatible
Model deepseek-v4-flash, deepseek-v4-pro
Base URL https://api.deepseek.com/v1
OPENAI_MAX_TOKENS 262144

Error Message

API Error: 400 {"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}

DeepSeek's Documented Rule

"When thinking: { type: 'enabled' } is active, all role=assistant messages in the conversation history must include reasoning_content."

— api-docs.deepseek.com/guides/thinking_mode

Based on the docs and the API behavior we observed, the practical rule seems to be:

  • Assistant message has a thinking block → echo the exact reasoning_content.
  • Assistant message does NOT have a thinking block → send reasoning_content: "" (empty string).
  • Assistant message omits the property entirely (undefined) → 400 in the DeepSeek scenarios reproduced here.

Validation: 5 CURLs Against Real DeepSeek API

CURL 1 — First request (user asks for tool use)

curl https://api.deepseek.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $DEEPSEEK_API_KEY"   -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "What time is it? Use the get_current_time tool."}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_current_time",
          "description": "Get current time",
          "parameters": {"type": "object", "properties": {}}
        }
      }
    ]
  }'

Response (200 OK):

{
    "id": "eeae50f5-f00c-4556-a999-200239f6ad0d",
    "object": "chat.completion",
    "created": 1777007567,
    "model": "deepseek-v4-flash",
    "choices": [{
        "index": 0,
        "message": {
            "role": "assistant",
            "content": "",
            "reasoning_content": "The user wants to know the current time. Let me use the get_current_time tool.",
            "tool_calls": [{
                "index": 0,
                "id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
                "type": "function",
                "function": {
                    "name": "get_current_time",
                    "arguments": "{}"
                }
            }]
        },
        "finish_reason": "tool_calls"
    }],
    "usage": {
        "prompt_tokens": 267,
        "completion_tokens": 47,
        "total_tokens": 314,
        "prompt_tokens_details": {"cached_tokens": 256},
        "completion_tokens_details": {"reasoning_tokens": 18},
        "prompt_cache_hit_tokens": 256,
        "prompt_cache_miss_tokens": 11
    }
}

DeepSeek returns content: "", reasoning_content, and tool_calls — the standard agent format. reasoning_tokens: 18 in usage suggests that thinking mode is active.


CURL 2 — Second request WITH reasoning_content → 200 OK

curl https://api.deepseek.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $DEEPSEEK_API_KEY"   -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "What time is it? Use the get_current_time tool."},
      {
        "role": "assistant",
        "content": "Let me check the current time for you.",
        "reasoning_content": "The user wants to know the current time. Let me use the get_current_time tool.",
        "tool_calls": [
          {
            "id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
            "type": "function",
            "function": {
              "name": "get_current_time",
              "arguments": "{}"
            }
          }
        ]
      },
      {"role": "tool", "tool_call_id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS", "content": "10:30 AM UTC"},
      {"role": "user", "content": "Thanks! What about tomorrow?"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_current_time",
          "description": "Get current time",
          "parameters": {"type": "object", "properties": {}}
        }
      }
    ]
  }'

Response (200 OK): Validated — conversation continues normally.


CURL 3 — Second request WITHOUT reasoning_content → 400 ERROR

curl https://api.deepseek.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $DEEPSEEK_API_KEY"   -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "What time is it? Use the get_current_time tool."},
      {
        "role": "assistant",
        "content": "Let me check the current time for you.",
        "tool_calls": [
          {
            "id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
            "type": "function",
            "function": {
              "name": "get_current_time",
              "arguments": "{}"
            }
          }
        ]
      },
      {"role": "tool", "tool_call_id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS", "content": "10:30 AM UTC"},
      {"role": "user", "content": "Thanks! What about tomorrow?"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_current_time",
          "description": "Get current time",
          "parameters": {"type": "object", "properties": {}}
        }
      }
    ]
  }'

Response (400 ERROR):

{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}

In this reproduction, the relevant difference from CURL 2 is the absence of reasoning_content on the assistant message.


CURL 4 — Agent format content: "" + tool_calls WITHOUT reasoning → 400

Exact format used by OpenClaude agent flows:

curl https://api.deepseek.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $DEEPSEEK_API_KEY"   -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Run ls"},
      {"role": "assistant", "content": "", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "Bash", "arguments": "{"command": "ls"}"}}]},
      {"role": "tool", "tool_call_id": "call_1", "content": "file.txt"}
    ],
    "stream": false
  }'

Response (400 ERROR):

{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}

CURL 5 — Agent format content: "" + tool_calls WITH reasoning → 200 OK

curl https://api.deepseek.com/v1/chat/completions   -H "Content-Type: application/json"   -H "Authorization: Bearer $DEEPSEEK_API_KEY"   -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Run ls"},
      {"role": "assistant", "content": "", "reasoning_content": "The user wants me to run ls. I will call the Bash tool.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "Bash", "arguments": "{"command": "ls"}"}}]},
      {"role": "tool", "tool_call_id": "call_1", "content": "file.txt"}
    ],
    "stream": false
  }'

Response: 200 OK


Validation Summary

# Format content reasoning_content Result
1 First request (user only) N/A N/A 200 OK
2 Multi-turn WITH reasoning "Let me check..." ✅ Present 200 OK
3 Multi-turn WITHOUT reasoning "Let me check..." ❌ Absent 400 ERROR
4 Agent format WITHOUT reasoning "" ❌ Absent 400 ERROR
5 Agent format WITH reasoning "" ✅ Present 200 OK

What Appear to Be the Four Edge Cases Not Covered by ff2a380

Important: These edge cases only appear to matter when thinking mode is enabled (thinking: { type: 'enabled' }). When thinking mode is OFF, DeepSeek does not seem to require reasoning_content, so these code paths may be irrelevant.

Commit ff2a380 appears to correctly handle the happy path: assistant messages that contain a thinking block. However, in real-world multi-turn conversations with thinking mode ON, we found four situations where reasoning_content was not being set:

Edge Case 1: Assistant with array content but NO thinking block

When the model calls tools without emitting a visible thinking block, the hasThinkingBlock guard prevented reasoning_content from being set at all.

// BEFORE (ff2a380, suspected bug):
if (preserveReasoningContent) {
  const thinkingText = (thinkingBlock as { thinking?: string })?.thinking
  if (typeof thinkingText === 'string' && thinkingText.trim().length > 0) {
    assistantMsg.reasoning_content = thinkingText  // <-- only if thinking exists
  }
}

// AFTER (proposed fix):
const reasoningContent = preserveReasoningContent
  ? ((thinkingBlock)?.thinking ?? '')  // <-- empty string if no thinking block
  : undefined
const shouldAttachReasoning = preserveReasoningContent  // <-- always for DeepSeek

Edge Case 2: Assistant with content as a string (not an array)

When content is a plain string (not an array of blocks), the else branch in convertMessages() created the assistant message without reasoning_content.

// AFTER (proposed fix):
} else {
  const assistantMsg = {
    role: 'assistant',
    content: ...,
    ...(preserveReasoningContent && {
      reasoning_content: '',
    }),
  }
}

Edge Case 3: Synthetic "[Tool execution interrupted by user]" message

The shim injects a synthetic assistant message when a tool message is followed by a user message (required for OpenAI/Mistral role alternation). This message was created as a bare object literal outside of convertMessages(), so it never received reasoning_content.

// BEFORE (suspected bug):
coalesced.push({
  role: 'assistant',
  content: '[Tool execution interrupted by user]',
})

// AFTER (proposed fix):
coalesced.push({
  role: 'assistant',
  content: '[Tool execution interrupted by user]',
  ...(preserveReasoningContent && {
    reasoning_content: '',
  }),
})

Edge Case 4: conversationRecovery.ts stripping thinking blocks

stripThinkingBlocks() removed all thinking/redacted_thinking blocks during conversation deserialization for 3P providers. This prevented the shim from ever seeing the thinking blocks to convert them to reasoning_content.

- function stripThinkingBlocks(messages) { ... }  // REMOVED
- const thinkingStripped = isThirdPartyProvider
-   ? stripThinkingBlocks(filteredThinking)
-   : filteredThinking
- const filteredMessages = filterWhitespaceOnlyAssistantMessages(thinkingStripped)
+ const filteredMessages = filterWhitespaceOnlyAssistantMessages(filteredThinking)

Why We Expect Limited Impact on Other Providers

All four changes are gated behind preserveReasoningContent, which is only true for DeepSeek and Moonshot:

// openaiShim.ts
preserveReasoningContent:
  isMoonshotBaseUrl(request.baseUrl) || isDeepSeekBaseUrl(request.baseUrl),

For providers where preserveReasoningContent remains false (OpenAI, Azure, Ollama, LM Studio, OpenRouter, Together, Groq, Fireworks, Mistral, Vertex, Bedrock, etc.), these changes are expected to be no-ops:

  • reasoningContent evaluates to undefined
  • shouldAttachReasoning is false
  • All ...(preserveReasoningContent && ...) spreads remain inactive

Honest uncertainty

Our validation was performed against DeepSeek. We have not exhaustively validated these paths end-to-end against other providers, so the non-DeepSeek impact assessment should be treated as a reasoned expectation rather than a fully verified claim.

A few specific uncertainties:

  • Moonshot / Kimi: Since preserveReasoningContent is also true there, these changes may affect it as well. We have not validated that behavior directly.
  • Future providers: Any provider later added to the preserveReasoningContent list would inherit these rules.

Files Modified

File Change
src/services/api/openaiShim.ts Edge cases 1, 2, 3: reasoning_content gating
src/utils/conversationRecovery.ts Edge case 4: removed stripThinkingBlocks()
src/utils/conversationRecovery.hooks.test.ts Updated test to reflect preserved thinking blocks

Why Subagents Are Most Affected

Agent flows in OpenClaude often use a pattern where many turns involve tool calls:

Turn 1: user -> assistant (+ reasoning_1 + tool_call_1) -> tool result
Turn 2: user -> assistant (+ reasoning_2 + tool_call_2) -> tool result
Turn 3: user -> assistant (+ reasoning_3 + tool_call_3) -> tool result
...

Each assistant message with tool_calls can carry reasoning_content that must be preserved. In a long agent session, the history may contain many assistant messages that need their reasoning_content echoed back. The four edge cases described above appear more likely to show up as the conversation grows.

Simple chat without tools works fine because no tool_calls means no reasoning_content requirement.


Request for Review

We are not core contributors and our understanding of the codebase is limited. We would greatly appreciate review from:

  • Jatmn — author of the original DeepSeek support (ff2a380) @jatmn
  • Anyone familiar with the OpenAI shim architecture
  • Anyone with access to test Moonshot / Kimi (the other provider currently using preserveReasoningContent)

Please help verify:

  1. Whether the gating logic is sufficient to protect other providers.
  2. Whether sending reasoning_content: "" (empty string) is the right fallback for assistant messages that never had thinking content.
  3. Whether removing stripThinkingBlocks() could have unintended effects on Moonshot or other 3P provider recovery paths.
  4. Whether there are cleaner places in the codebase to address these edge cases.

Related Issues


References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions