[Bug]: DeepSeek V4 — 400 error caused by reasoning_content missing on edge-case assistant messages
Author's note: This report builds on the foundational DeepSeek V4 support added by Jatmn in commit ff2a380 (Add DeepSeek V4 flash/pro support and DeepSeek thinking compatibility #877). That commit appears to have provided the core infrastructure: DeepSeek URL detection, thinking: { type: 'enabled' } support, and the preserveReasoningContent mechanism. We are very grateful for that work.
What we found: After testing extensively with deepseek-v4-pro, we found what appear to be four edge cases not covered by ff2a380 where reasoning_content was not being set, which in our DeepSeek reproductions led to 400 errors in multi-turn conversations with tool use. The original commit seems to handle the happy path well (assistant messages that do contain thinking blocks), but may miss assistant messages without thinking blocks, string-content messages, and synthetic interrupt messages.
Our validation for this report was performed against DeepSeek. While the proposed explanation and impact assessment seem reasonable, we are not fully certain there are no secondary effects in other paths or providers, so review from maintainers with deeper context would be greatly appreciated.
Summary
When using OpenClaude with DeepSeek V4 (deepseek-v4-flash, deepseek-v4-pro) as an OpenAI-compatible provider, conversations that involve tool calls can fail with a 400 invalid_request_error. The error message is:
The `reasoning_content` in the thinking mode must be passed back to the API.
This appears to happen because:
- DeepSeek V4 enables thinking mode by default, which emits
reasoning_content on assistant messages.
- Commit
ff2a380 added the infrastructure to capture thinking blocks and re-attach them as reasoning_content.
- However, what appear to be four edge cases were not covered, causing
reasoning_content to be undefined on some assistant messages.
- In the DeepSeek scenarios we validated, the request is rejected if an assistant message in thinking mode omits
reasoning_content.
This issue appears to affect tool-calling conversations with DeepSeek V4, especially subagents / Task tool invocations where tools are used repeatedly across turns.
Environment
| Component |
Version / Setting |
| OpenClaude |
v0.6.0 |
| Provider |
openai-compatible |
| Model |
deepseek-v4-flash, deepseek-v4-pro |
| Base URL |
https://api.deepseek.com/v1 |
OPENAI_MAX_TOKENS |
262144 |
Error Message
API Error: 400 {"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}
DeepSeek's Documented Rule
"When thinking: { type: 'enabled' } is active, all role=assistant messages in the conversation history must include reasoning_content."
— api-docs.deepseek.com/guides/thinking_mode
Based on the docs and the API behavior we observed, the practical rule seems to be:
- Assistant message has a thinking block → echo the exact
reasoning_content.
- Assistant message does NOT have a thinking block → send
reasoning_content: "" (empty string).
- Assistant message omits the property entirely (
undefined) → 400 in the DeepSeek scenarios reproduced here.
Validation: 5 CURLs Against Real DeepSeek API
CURL 1 — First request (user asks for tool use)
curl https://api.deepseek.com/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $DEEPSEEK_API_KEY" -d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "What time is it? Use the get_current_time tool."}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get current time",
"parameters": {"type": "object", "properties": {}}
}
}
]
}'
Response (200 OK):
{
"id": "eeae50f5-f00c-4556-a999-200239f6ad0d",
"object": "chat.completion",
"created": 1777007567,
"model": "deepseek-v4-flash",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "",
"reasoning_content": "The user wants to know the current time. Let me use the get_current_time tool.",
"tool_calls": [{
"index": 0,
"id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
"type": "function",
"function": {
"name": "get_current_time",
"arguments": "{}"
}
}]
},
"finish_reason": "tool_calls"
}],
"usage": {
"prompt_tokens": 267,
"completion_tokens": 47,
"total_tokens": 314,
"prompt_tokens_details": {"cached_tokens": 256},
"completion_tokens_details": {"reasoning_tokens": 18},
"prompt_cache_hit_tokens": 256,
"prompt_cache_miss_tokens": 11
}
}
DeepSeek returns content: "", reasoning_content, and tool_calls — the standard agent format. reasoning_tokens: 18 in usage suggests that thinking mode is active.
CURL 2 — Second request WITH reasoning_content → 200 OK
curl https://api.deepseek.com/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $DEEPSEEK_API_KEY" -d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "What time is it? Use the get_current_time tool."},
{
"role": "assistant",
"content": "Let me check the current time for you.",
"reasoning_content": "The user wants to know the current time. Let me use the get_current_time tool.",
"tool_calls": [
{
"id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
"type": "function",
"function": {
"name": "get_current_time",
"arguments": "{}"
}
}
]
},
{"role": "tool", "tool_call_id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS", "content": "10:30 AM UTC"},
{"role": "user", "content": "Thanks! What about tomorrow?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get current time",
"parameters": {"type": "object", "properties": {}}
}
}
]
}'
Response (200 OK): Validated — conversation continues normally.
CURL 3 — Second request WITHOUT reasoning_content → 400 ERROR
curl https://api.deepseek.com/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $DEEPSEEK_API_KEY" -d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "What time is it? Use the get_current_time tool."},
{
"role": "assistant",
"content": "Let me check the current time for you.",
"tool_calls": [
{
"id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS",
"type": "function",
"function": {
"name": "get_current_time",
"arguments": "{}"
}
}
]
},
{"role": "tool", "tool_call_id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS", "content": "10:30 AM UTC"},
{"role": "user", "content": "Thanks! What about tomorrow?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get current time",
"parameters": {"type": "object", "properties": {}}
}
}
]
}'
Response (400 ERROR):
{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}
In this reproduction, the relevant difference from CURL 2 is the absence of reasoning_content on the assistant message.
CURL 4 — Agent format content: "" + tool_calls WITHOUT reasoning → 400
Exact format used by OpenClaude agent flows:
curl https://api.deepseek.com/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $DEEPSEEK_API_KEY" -d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Run ls"},
{"role": "assistant", "content": "", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "Bash", "arguments": "{"command": "ls"}"}}]},
{"role": "tool", "tool_call_id": "call_1", "content": "file.txt"}
],
"stream": false
}'
Response (400 ERROR):
{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}
CURL 5 — Agent format content: "" + tool_calls WITH reasoning → 200 OK
curl https://api.deepseek.com/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $DEEPSEEK_API_KEY" -d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Run ls"},
{"role": "assistant", "content": "", "reasoning_content": "The user wants me to run ls. I will call the Bash tool.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "Bash", "arguments": "{"command": "ls"}"}}]},
{"role": "tool", "tool_call_id": "call_1", "content": "file.txt"}
],
"stream": false
}'
Response: 200 OK
Validation Summary
| # |
Format |
content |
reasoning_content |
Result |
| 1 |
First request (user only) |
N/A |
N/A |
200 OK |
| 2 |
Multi-turn WITH reasoning |
"Let me check..." |
✅ Present |
200 OK |
| 3 |
Multi-turn WITHOUT reasoning |
"Let me check..." |
❌ Absent |
400 ERROR |
| 4 |
Agent format WITHOUT reasoning |
"" |
❌ Absent |
400 ERROR |
| 5 |
Agent format WITH reasoning |
"" |
✅ Present |
200 OK |
What Appear to Be the Four Edge Cases Not Covered by ff2a380
Important: These edge cases only appear to matter when thinking mode is enabled (thinking: { type: 'enabled' }). When thinking mode is OFF, DeepSeek does not seem to require reasoning_content, so these code paths may be irrelevant.
Commit ff2a380 appears to correctly handle the happy path: assistant messages that contain a thinking block. However, in real-world multi-turn conversations with thinking mode ON, we found four situations where reasoning_content was not being set:
Edge Case 1: Assistant with array content but NO thinking block
When the model calls tools without emitting a visible thinking block, the hasThinkingBlock guard prevented reasoning_content from being set at all.
// BEFORE (ff2a380, suspected bug):
if (preserveReasoningContent) {
const thinkingText = (thinkingBlock as { thinking?: string })?.thinking
if (typeof thinkingText === 'string' && thinkingText.trim().length > 0) {
assistantMsg.reasoning_content = thinkingText // <-- only if thinking exists
}
}
// AFTER (proposed fix):
const reasoningContent = preserveReasoningContent
? ((thinkingBlock)?.thinking ?? '') // <-- empty string if no thinking block
: undefined
const shouldAttachReasoning = preserveReasoningContent // <-- always for DeepSeek
Edge Case 2: Assistant with content as a string (not an array)
When content is a plain string (not an array of blocks), the else branch in convertMessages() created the assistant message without reasoning_content.
// AFTER (proposed fix):
} else {
const assistantMsg = {
role: 'assistant',
content: ...,
...(preserveReasoningContent && {
reasoning_content: '',
}),
}
}
Edge Case 3: Synthetic "[Tool execution interrupted by user]" message
The shim injects a synthetic assistant message when a tool message is followed by a user message (required for OpenAI/Mistral role alternation). This message was created as a bare object literal outside of convertMessages(), so it never received reasoning_content.
// BEFORE (suspected bug):
coalesced.push({
role: 'assistant',
content: '[Tool execution interrupted by user]',
})
// AFTER (proposed fix):
coalesced.push({
role: 'assistant',
content: '[Tool execution interrupted by user]',
...(preserveReasoningContent && {
reasoning_content: '',
}),
})
Edge Case 4: conversationRecovery.ts stripping thinking blocks
stripThinkingBlocks() removed all thinking/redacted_thinking blocks during conversation deserialization for 3P providers. This prevented the shim from ever seeing the thinking blocks to convert them to reasoning_content.
- function stripThinkingBlocks(messages) { ... } // REMOVED
- const thinkingStripped = isThirdPartyProvider
- ? stripThinkingBlocks(filteredThinking)
- : filteredThinking
- const filteredMessages = filterWhitespaceOnlyAssistantMessages(thinkingStripped)
+ const filteredMessages = filterWhitespaceOnlyAssistantMessages(filteredThinking)
Why We Expect Limited Impact on Other Providers
All four changes are gated behind preserveReasoningContent, which is only true for DeepSeek and Moonshot:
// openaiShim.ts
preserveReasoningContent:
isMoonshotBaseUrl(request.baseUrl) || isDeepSeekBaseUrl(request.baseUrl),
For providers where preserveReasoningContent remains false (OpenAI, Azure, Ollama, LM Studio, OpenRouter, Together, Groq, Fireworks, Mistral, Vertex, Bedrock, etc.), these changes are expected to be no-ops:
reasoningContent evaluates to undefined
shouldAttachReasoning is false
- All
...(preserveReasoningContent && ...) spreads remain inactive
Honest uncertainty
Our validation was performed against DeepSeek. We have not exhaustively validated these paths end-to-end against other providers, so the non-DeepSeek impact assessment should be treated as a reasoned expectation rather than a fully verified claim.
A few specific uncertainties:
- Moonshot / Kimi: Since
preserveReasoningContent is also true there, these changes may affect it as well. We have not validated that behavior directly.
- Future providers: Any provider later added to the
preserveReasoningContent list would inherit these rules.
Files Modified
| File |
Change |
src/services/api/openaiShim.ts |
Edge cases 1, 2, 3: reasoning_content gating |
src/utils/conversationRecovery.ts |
Edge case 4: removed stripThinkingBlocks() |
src/utils/conversationRecovery.hooks.test.ts |
Updated test to reflect preserved thinking blocks |
Why Subagents Are Most Affected
Agent flows in OpenClaude often use a pattern where many turns involve tool calls:
Turn 1: user -> assistant (+ reasoning_1 + tool_call_1) -> tool result
Turn 2: user -> assistant (+ reasoning_2 + tool_call_2) -> tool result
Turn 3: user -> assistant (+ reasoning_3 + tool_call_3) -> tool result
...
Each assistant message with tool_calls can carry reasoning_content that must be preserved. In a long agent session, the history may contain many assistant messages that need their reasoning_content echoed back. The four edge cases described above appear more likely to show up as the conversation grows.
Simple chat without tools works fine because no tool_calls means no reasoning_content requirement.
Request for Review
We are not core contributors and our understanding of the codebase is limited. We would greatly appreciate review from:
- Jatmn — author of the original DeepSeek support (
ff2a380) @jatmn
- Anyone familiar with the OpenAI shim architecture
- Anyone with access to test Moonshot / Kimi (the other provider currently using
preserveReasoningContent)
Please help verify:
- Whether the gating logic is sufficient to protect other providers.
- Whether sending
reasoning_content: "" (empty string) is the right fallback for assistant messages that never had thinking content.
- Whether removing
stripThinkingBlocks() could have unintended effects on Moonshot or other 3P provider recovery paths.
- Whether there are cleaner places in the codebase to address these edge cases.
Related Issues
References
[Bug]: DeepSeek V4 — 400 error caused by
reasoning_contentmissing on edge-case assistant messagesSummary
When using OpenClaude with DeepSeek V4 (
deepseek-v4-flash,deepseek-v4-pro) as an OpenAI-compatible provider, conversations that involve tool calls can fail with a400 invalid_request_error. The error message is:This appears to happen because:
reasoning_contenton assistant messages.ff2a380added the infrastructure to capturethinkingblocks and re-attach them asreasoning_content.reasoning_contentto beundefinedon some assistant messages.reasoning_content.This issue appears to affect tool-calling conversations with DeepSeek V4, especially subagents / Task tool invocations where tools are used repeatedly across turns.
Environment
v0.6.0openai-compatibledeepseek-v4-flash,deepseek-v4-prohttps://api.deepseek.com/v1OPENAI_MAX_TOKENS262144Error Message
DeepSeek's Documented Rule
Based on the docs and the API behavior we observed, the practical rule seems to be:
reasoning_content.reasoning_content: ""(empty string).undefined) → 400 in the DeepSeek scenarios reproduced here.Validation: 5 CURLs Against Real DeepSeek API
CURL 1 — First request (user asks for tool use)
Response (200 OK):
{ "id": "eeae50f5-f00c-4556-a999-200239f6ad0d", "object": "chat.completion", "created": 1777007567, "model": "deepseek-v4-flash", "choices": [{ "index": 0, "message": { "role": "assistant", "content": "", "reasoning_content": "The user wants to know the current time. Let me use the get_current_time tool.", "tool_calls": [{ "index": 0, "id": "call_00_9nLfpIQ6VO1JCaEsJcU8WZnS", "type": "function", "function": { "name": "get_current_time", "arguments": "{}" } }] }, "finish_reason": "tool_calls" }], "usage": { "prompt_tokens": 267, "completion_tokens": 47, "total_tokens": 314, "prompt_tokens_details": {"cached_tokens": 256}, "completion_tokens_details": {"reasoning_tokens": 18}, "prompt_cache_hit_tokens": 256, "prompt_cache_miss_tokens": 11 } }DeepSeek returns
content: "",reasoning_content, andtool_calls— the standard agent format.reasoning_tokens: 18in usage suggests that thinking mode is active.CURL 2 — Second request WITH
reasoning_content→ 200 OKResponse (200 OK): Validated — conversation continues normally.
CURL 3 — Second request WITHOUT
reasoning_content→ 400 ERRORResponse (400 ERROR):
{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}In this reproduction, the relevant difference from CURL 2 is the absence of
reasoning_contenton the assistant message.CURL 4 — Agent format
content: ""+tool_callsWITHOUT reasoning → 400Exact format used by OpenClaude agent flows:
Response (400 ERROR):
{"error":{"message":"The `reasoning_content` in the thinking mode must be passed back to the API.","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}CURL 5 — Agent format
content: ""+tool_callsWITH reasoning → 200 OKResponse: 200 OK
Validation Summary
contentreasoning_content"Let me check...""Let me check..."""""What Appear to Be the Four Edge Cases Not Covered by
ff2a380Commit
ff2a380appears to correctly handle the happy path: assistant messages that contain athinkingblock. However, in real-world multi-turn conversations with thinking mode ON, we found four situations wherereasoning_contentwas not being set:Edge Case 1: Assistant with array
contentbut NO thinking blockWhen the model calls tools without emitting a visible thinking block, the
hasThinkingBlockguard preventedreasoning_contentfrom being set at all.Edge Case 2: Assistant with
contentas a string (not an array)When
contentis a plain string (not an array of blocks), theelsebranch inconvertMessages()created the assistant message withoutreasoning_content.Edge Case 3: Synthetic
"[Tool execution interrupted by user]"messageThe shim injects a synthetic assistant message when a
toolmessage is followed by ausermessage (required for OpenAI/Mistral role alternation). This message was created as a bare object literal outside ofconvertMessages(), so it never receivedreasoning_content.Edge Case 4:
conversationRecovery.tsstripping thinking blocksstripThinkingBlocks()removed allthinking/redacted_thinkingblocks during conversation deserialization for 3P providers. This prevented the shim from ever seeing the thinking blocks to convert them toreasoning_content.Why We Expect Limited Impact on Other Providers
All four changes are gated behind
preserveReasoningContent, which is onlytruefor DeepSeek and Moonshot:For providers where
preserveReasoningContentremainsfalse(OpenAI, Azure, Ollama, LM Studio, OpenRouter, Together, Groq, Fireworks, Mistral, Vertex, Bedrock, etc.), these changes are expected to be no-ops:reasoningContentevaluates toundefinedshouldAttachReasoningisfalse...(preserveReasoningContent && ...)spreads remain inactiveHonest uncertainty
Our validation was performed against DeepSeek. We have not exhaustively validated these paths end-to-end against other providers, so the non-DeepSeek impact assessment should be treated as a reasoned expectation rather than a fully verified claim.
A few specific uncertainties:
preserveReasoningContentis alsotruethere, these changes may affect it as well. We have not validated that behavior directly.preserveReasoningContentlist would inherit these rules.Files Modified
src/services/api/openaiShim.tsreasoning_contentgatingsrc/utils/conversationRecovery.tsstripThinkingBlocks()src/utils/conversationRecovery.hooks.test.tsWhy Subagents Are Most Affected
Agent flows in OpenClaude often use a pattern where many turns involve tool calls:
Each assistant message with
tool_callscan carryreasoning_contentthat must be preserved. In a long agent session, the history may contain many assistant messages that need theirreasoning_contentechoed back. The four edge cases described above appear more likely to show up as the conversation grows.Simple chat without tools works fine because no
tool_callsmeans noreasoning_contentrequirement.Request for Review
We are not core contributors and our understanding of the codebase is limited. We would greatly appreciate review from:
ff2a380) @jatmnpreserveReasoningContent)Please help verify:
reasoning_content: ""(empty string) is the right fallback for assistant messages that never had thinking content.stripThinkingBlocks()could have unintended effects on Moonshot or other 3P provider recovery paths.Related Issues
reasoning_content missing problem(Moonshot — fixed by commit67de6bd2)References
ff2a3807230101b6e170b5ecea8f6aef7a90ece9—Add DeepSeek V4 flash/pro support and DeepSeek thinking compatibility (#877)by Jatmn67de6bd2cffc3381f0f28fd3ffce043970611667—fix(openai-shim): echo reasoning_content on assistant tool-call messages for Moonshot (#828)