Description
Anthropic reports overloads differently for non-streaming and streaming requests. Its HTTP error documentation assigns HTTP 529 to a non-streaming overloaded_error. Its streaming documentation explains that a streaming request can return HTTP 200 and report the same condition later through the following Server-Sent Events (SSE) event:
HTTP/1.1 200 OK
Content-Type: text/event-stream
event: error
data: {"type":"error","error":{"type":"overloaded_error","message":"Overloaded"}}
The current Chat Completions translator parses the Anthropic event but returns a Go error instead of a translated stream item:
case string(constant.ValueOf[constant.Error]()):
var errEvent anthropic.ErrorResponse
if err := json.Unmarshal(data, &errEvent); err != nil {
return nil, fmt.Errorf("unparsable error event: %s", string(data))
}
return nil, fmt.Errorf("anthropic stream error: %s - %s", errEvent.Error.Type, errEvent.Error.Message)
The response processor propagates that translator error instead of returning a response-body mutation or an immediate response:
newHeaders, newBody, tokenUsage, responseModel, err := u.translator.ResponseBody(
u.responseHeaders, decodingResult.reader, body.EndOfStream, u.parent.span,
)
if err != nil {
return nil, fmt.Errorf("failed to transform response: %w", err)
}
The propagated Go error closes the external-processing remote procedure call (RPC). Under Envoy's external processor failure behavior, what the client receives depends on whether Envoy has committed the downstream response:
| Downstream state |
Envoy behavior |
Client-visible result |
| HTTP response not committed |
Envoy replaces the upstream HTTP 200 response with a local HTTP 500 response. |
The request fails with a generic HTTP 500 before stream iteration begins. |
| HTTP 200 response committed |
Envoy cannot change the status, so it terminates the existing stream. |
The client receives HTTP 200, but stream iteration later fails because the stream is truncated. |
Neither result exposes the provider's overloaded_error classification, so users cannot distinguish an overload from a generic gateway failure. The gateway needs a downstream representation that preserves this distinction for Chat Completions clients.
Option 1: synthesize HTTP 529
If Envoy has not committed the downstream response, the gateway can replace the upstream HTTP 200 stream with a non-streaming error response:
HTTP/1.1 529
Content-Type: application/json
{"error":{"type":"overloaded_error","message":"Overloaded"}}
This puts the overload on the HTTP error path. Clients receive the failure before stream iteration begins, status-based monitoring records it, and HTTP 5xx retry policies may treat it as retryable. However, a 529-only implementation still has two possible results:
| Downstream state when overload arrives |
529 result |
| HTTP response not committed |
Return HTTP 529 with a JSON error. |
| HTTP 200 response committed |
Terminate the stream because the status can no longer be changed. |
The limitation is inconsistency: the same provider error has two client-visible outcomes because an HTTP status cannot be changed after the response is committed. Returning 529 also differs from Anthropic's documented streaming behavior, which reports the error in-band under HTTP 200.
Option 2: preserve HTTP 200 and emit an in-band error
The gateway can preserve the upstream streaming status and translate the Anthropic event into a data-only top-level error envelope:
HTTP/1.1 200 OK
Content-Type: text/event-stream
data: {"error":{"type":"overloaded_error","message":"Overloaded"}}
This option behaves consistently for every streaming overload: it preserves HTTP 200 and sends the same terminal error event downstream. It follows Anthropic's documented streaming model—HTTP 200 followed by an in-band error—while translating the envelope for Chat Completions clients.
The compatibility limitation is that the Chat Completions streaming reference documents chat.completion.chunk events but does not define a terminal top-level error event. The proposed event is handled by the SDKs' streaming layers rather than parsed as a ChatCompletionChunk.
The official OpenAI Python v2.13.0, Node v7.10.0, and Go v3.41.0 stream parsers check for a top-level error before parsing a normal stream item. The Python SDK raises APIError during iteration and exposes the nested object as error.body.
A structural review of gateway response logs also found that direct OpenAI Chat Completions streams used a data-only SSE item containing a top-level error object and no named event: line:
data: {"error":{"code":"internal_server_error","message":"...","param":null,"type":"server_error"}}
The cited SDK versions therefore handle the same top-level envelope observed in direct OpenAI Chat Completions traffic. Python clients can classify the proposed overload event by checking error.body["type"] == "overloaded_error" instead of parsing exception text. However, the event is not a documented ChatCompletionChunk variant, and support cannot be assumed for every third-party Chat Completions client.
Proposed resolution
Use HTTP 200 with a terminal in-band error for every streaming overload. Of the two options, this is the only one that preserves overloaded_error after an HTTP 200 stream has started. It also avoids changing client behavior based on response timing. The event should contain only the overload type and message; provider request IDs and diagnostic details must not be copied.
The gateway should:
- Return a typed translator error for the exact Anthropic
overloaded_error type while retaining the serialized top-level error event.
- Return that event as a normal response-body mutation under HTTP 200.
- Mark the stream terminal and suppress every later provider body chunk.
- Record the request as failed in metrics and tracing, and omit
[DONE].
- Continue propagating every other Anthropic streaming error through the existing failure path.
Acceptance criteria
- An overload as the first provider SSE event returns HTTP 200 and
data: {"error":{"type":"overloaded_error","message":"Overloaded"}}.
- An overload after one or more translated chunks returns the same error event after those chunks.
- Content followed by overload in the same response-body callback preserves the content and ends with the error event.
- No later provider event and no
[DONE] appears after the overload.
- The request is recorded as failed in metrics and tracing.
- Non-overload Anthropic SSE errors retain their current behavior.
- Tests exercise the error as the first event, after content, in the same callback as content, and with later provider chunks.
- Compatibility is verified with supported OpenAI client versions by checking the structured error body during iteration.
Description
Anthropic reports overloads differently for non-streaming and streaming requests. Its HTTP error documentation assigns HTTP 529 to a non-streaming
overloaded_error. Its streaming documentation explains that a streaming request can return HTTP 200 and report the same condition later through the following Server-Sent Events (SSE) event:The current Chat Completions translator parses the Anthropic event but returns a Go error instead of a translated stream item:
The response processor propagates that translator error instead of returning a response-body mutation or an immediate response:
The propagated Go error closes the external-processing remote procedure call (RPC). Under Envoy's external processor failure behavior, what the client receives depends on whether Envoy has committed the downstream response:
Neither result exposes the provider's
overloaded_errorclassification, so users cannot distinguish an overload from a generic gateway failure. The gateway needs a downstream representation that preserves this distinction for Chat Completions clients.Option 1: synthesize HTTP 529
If Envoy has not committed the downstream response, the gateway can replace the upstream HTTP 200 stream with a non-streaming error response:
This puts the overload on the HTTP error path. Clients receive the failure before stream iteration begins, status-based monitoring records it, and HTTP 5xx retry policies may treat it as retryable. However, a 529-only implementation still has two possible results:
The limitation is inconsistency: the same provider error has two client-visible outcomes because an HTTP status cannot be changed after the response is committed. Returning 529 also differs from Anthropic's documented streaming behavior, which reports the error in-band under HTTP 200.
Option 2: preserve HTTP 200 and emit an in-band error
The gateway can preserve the upstream streaming status and translate the Anthropic event into a data-only top-level error envelope:
This option behaves consistently for every streaming overload: it preserves HTTP 200 and sends the same terminal error event downstream. It follows Anthropic's documented streaming model—HTTP 200 followed by an in-band error—while translating the envelope for Chat Completions clients.
The compatibility limitation is that the Chat Completions streaming reference documents
chat.completion.chunkevents but does not define a terminal top-levelerrorevent. The proposed event is handled by the SDKs' streaming layers rather than parsed as aChatCompletionChunk.The official OpenAI Python v2.13.0, Node v7.10.0, and Go v3.41.0 stream parsers check for a top-level
errorbefore parsing a normal stream item. The Python SDK raisesAPIErrorduring iteration and exposes the nested object aserror.body.A structural review of gateway response logs also found that direct OpenAI Chat Completions streams used a data-only SSE item containing a top-level
errorobject and no namedevent:line:The cited SDK versions therefore handle the same top-level envelope observed in direct OpenAI Chat Completions traffic. Python clients can classify the proposed overload event by checking
error.body["type"] == "overloaded_error"instead of parsing exception text. However, the event is not a documentedChatCompletionChunkvariant, and support cannot be assumed for every third-party Chat Completions client.Proposed resolution
Use HTTP 200 with a terminal in-band error for every streaming overload. Of the two options, this is the only one that preserves
overloaded_errorafter an HTTP 200 stream has started. It also avoids changing client behavior based on response timing. The event should contain only the overload type and message; provider request IDs and diagnostic details must not be copied.The gateway should:
overloaded_errortype while retaining the serialized top-level error event.[DONE].Acceptance criteria
data: {"error":{"type":"overloaded_error","message":"Overloaded"}}.[DONE]appears after the overload.