Skip to content

Preserve Anthropic streaming overload errors for Chat Completions clients #2697

Description

@hustxiayang

Description

Anthropic reports overloads differently for non-streaming and streaming requests. Its HTTP error documentation assigns HTTP 529 to a non-streaming overloaded_error. Its streaming documentation explains that a streaming request can return HTTP 200 and report the same condition later through the following Server-Sent Events (SSE) event:

HTTP/1.1 200 OK
Content-Type: text/event-stream

event: error
data: {"type":"error","error":{"type":"overloaded_error","message":"Overloaded"}}

The current Chat Completions translator parses the Anthropic event but returns a Go error instead of a translated stream item:

case string(constant.ValueOf[constant.Error]()):
	var errEvent anthropic.ErrorResponse
	if err := json.Unmarshal(data, &errEvent); err != nil {
		return nil, fmt.Errorf("unparsable error event: %s", string(data))
	}
	return nil, fmt.Errorf("anthropic stream error: %s - %s", errEvent.Error.Type, errEvent.Error.Message)

The response processor propagates that translator error instead of returning a response-body mutation or an immediate response:

newHeaders, newBody, tokenUsage, responseModel, err := u.translator.ResponseBody(
	u.responseHeaders, decodingResult.reader, body.EndOfStream, u.parent.span,
)
if err != nil {
	return nil, fmt.Errorf("failed to transform response: %w", err)
}

The propagated Go error closes the external-processing remote procedure call (RPC). Under Envoy's external processor failure behavior, what the client receives depends on whether Envoy has committed the downstream response:

Downstream state Envoy behavior Client-visible result
HTTP response not committed Envoy replaces the upstream HTTP 200 response with a local HTTP 500 response. The request fails with a generic HTTP 500 before stream iteration begins.
HTTP 200 response committed Envoy cannot change the status, so it terminates the existing stream. The client receives HTTP 200, but stream iteration later fails because the stream is truncated.

Neither result exposes the provider's overloaded_error classification, so users cannot distinguish an overload from a generic gateway failure. The gateway needs a downstream representation that preserves this distinction for Chat Completions clients.

Option 1: synthesize HTTP 529

If Envoy has not committed the downstream response, the gateway can replace the upstream HTTP 200 stream with a non-streaming error response:

HTTP/1.1 529
Content-Type: application/json

{"error":{"type":"overloaded_error","message":"Overloaded"}}

This puts the overload on the HTTP error path. Clients receive the failure before stream iteration begins, status-based monitoring records it, and HTTP 5xx retry policies may treat it as retryable. However, a 529-only implementation still has two possible results:

Downstream state when overload arrives 529 result
HTTP response not committed Return HTTP 529 with a JSON error.
HTTP 200 response committed Terminate the stream because the status can no longer be changed.

The limitation is inconsistency: the same provider error has two client-visible outcomes because an HTTP status cannot be changed after the response is committed. Returning 529 also differs from Anthropic's documented streaming behavior, which reports the error in-band under HTTP 200.

Option 2: preserve HTTP 200 and emit an in-band error

The gateway can preserve the upstream streaming status and translate the Anthropic event into a data-only top-level error envelope:

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: {"error":{"type":"overloaded_error","message":"Overloaded"}}

This option behaves consistently for every streaming overload: it preserves HTTP 200 and sends the same terminal error event downstream. It follows Anthropic's documented streaming model—HTTP 200 followed by an in-band error—while translating the envelope for Chat Completions clients.

The compatibility limitation is that the Chat Completions streaming reference documents chat.completion.chunk events but does not define a terminal top-level error event. The proposed event is handled by the SDKs' streaming layers rather than parsed as a ChatCompletionChunk.

The official OpenAI Python v2.13.0, Node v7.10.0, and Go v3.41.0 stream parsers check for a top-level error before parsing a normal stream item. The Python SDK raises APIError during iteration and exposes the nested object as error.body.

A structural review of gateway response logs also found that direct OpenAI Chat Completions streams used a data-only SSE item containing a top-level error object and no named event: line:

data: {"error":{"code":"internal_server_error","message":"...","param":null,"type":"server_error"}}

The cited SDK versions therefore handle the same top-level envelope observed in direct OpenAI Chat Completions traffic. Python clients can classify the proposed overload event by checking error.body["type"] == "overloaded_error" instead of parsing exception text. However, the event is not a documented ChatCompletionChunk variant, and support cannot be assumed for every third-party Chat Completions client.

Proposed resolution

Use HTTP 200 with a terminal in-band error for every streaming overload. Of the two options, this is the only one that preserves overloaded_error after an HTTP 200 stream has started. It also avoids changing client behavior based on response timing. The event should contain only the overload type and message; provider request IDs and diagnostic details must not be copied.

The gateway should:

  1. Return a typed translator error for the exact Anthropic overloaded_error type while retaining the serialized top-level error event.
  2. Return that event as a normal response-body mutation under HTTP 200.
  3. Mark the stream terminal and suppress every later provider body chunk.
  4. Record the request as failed in metrics and tracing, and omit [DONE].
  5. Continue propagating every other Anthropic streaming error through the existing failure path.

Acceptance criteria

  • An overload as the first provider SSE event returns HTTP 200 and data: {"error":{"type":"overloaded_error","message":"Overloaded"}}.
  • An overload after one or more translated chunks returns the same error event after those chunks.
  • Content followed by overload in the same response-body callback preserves the content and ends with the error event.
  • No later provider event and no [DONE] appears after the overload.
  • The request is recorded as failed in metrics and tracing.
  • Non-overload Anthropic SSE errors retain their current behavior.
  • Tests exercise the error as the first event, after content, in the same callback as content, and with later provider chunks.
  • Compatibility is verified with supported OpenAI client versions by checking the structured error body during iteration.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesttriageNeeds initial triage by a maintainer

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions