Description:
Envoy AI Gateway currently performs request translation in the per-cluster upstream AI Gateway ext-proc, while provider response translation is intentionally handled by the AI Gateway ext-proc in the main downstream filter chain. Upstream response processing is disabled to preserve the existing retry and fallback lifecycle.
With a user caching ext-proc placed after the downstream AI Gateway ext-proc in request order, the current flow is: Request:
Client
→ AI Gateway downstream ext-proc (Model extraction)
→ caching ext-proc
→ Envoy router
→ AI Gateway upstream ext-proc
→ LLM provider
Because Envoy executes response filters in reverse order, the response flow becomes:
Response:
LLM provider
→ AI Gateway upstream ext-proc, response processing skipped
→ Envoy router
→ caching ext-proc
→ AI Gateway downstream ext-proc translates the response
→ Client
This means the caching extension observes and stores the provider-native response before AI Gateway translates it into the canonical OpenAI format.
- On a cache miss, the client receives the translated OpenAI response.
- On the following cache hit, the caching extension returns the stored provider-native response through an Immediate Response, bypassing AI Gateway response translation.
As a result, cache misses and cache hits can return different response formats for providers
such as AWS Bedrock and GCP Vertex.
Moving response translation into the upstream AI Gateway ext-proc could affect the current retry, fallback, streaming, metrics, and token-accounting lifecycle. The preferred solution is therefore to preserve the existing translation behavior and allow the caching extension to be placed before the downstream AI Gateway ext-proc in request order.
The desired filter order is: Request:
Client
→ caching ext-proc
→ AI Gateway downstream ext-proc
→ Envoy router
→ AI Gateway upstream ext-proc
→ LLM provider
The reverse response order would then be:
Response:
LLM provider
→ AI Gateway upstream ext-proc, response processing skipped
→ Envoy router
→ AI Gateway downstream ext-proc translates the response
→ caching ext-proc stores the canonical OpenAI response
→ Client
This ordering allows the caching extension to hash the original canonical OpenAI request during request processing and observe the translated canonical OpenAI response during response processing. A cache hit can then return the stored response directly without requiring provider-specific translation.
The proposed AIGatewayExtensionPolicy appears to provide the required mechanism. It inserts a composite-wrapped user ext-proc before the downstream AI Gateway ext-proc and enables it on the relevant generated routes, including the initial catch-all route used before x-ai-eg-model is derived. Since Clear Route Cache does not restart filters that have already executed, the caching extension should run once during request processing and receive the response later through the reverse filter path.
To support response caching, the policy must allow response processing rather than using response: Skip. Streaming LLM responses should use streamed response- body processing so that the extension can forward translated SSE chunks while accumulating them for the eventual cache write.
[optional Relevant Links:]
https://github.com/siddharth1036/ai-gateway/blob/91bfb1679cd71fa2b17e0954b3e606becb8c44c8/docs/proposals/012-aigatewayroutefilter/proposal.md
#2364
#2671
Description:
Envoy AI Gateway currently performs request translation in the per-cluster upstream AI Gateway ext-proc, while provider response translation is intentionally handled by the AI Gateway ext-proc in the main downstream filter chain. Upstream response processing is disabled to preserve the existing retry and fallback lifecycle.
With a user caching ext-proc placed after the downstream AI Gateway ext-proc in request order, the current flow is: Request:
Client
→ AI Gateway downstream ext-proc (Model extraction)
→ caching ext-proc
→ Envoy router
→ AI Gateway upstream ext-proc
→ LLM provider
Because Envoy executes response filters in reverse order, the response flow becomes:
Response:
LLM provider
→ AI Gateway upstream ext-proc, response processing skipped
→ Envoy router
→ caching ext-proc
→ AI Gateway downstream ext-proc translates the response
→ Client
This means the caching extension observes and stores the provider-native response before AI Gateway translates it into the canonical OpenAI format.
As a result, cache misses and cache hits can return different response formats for providers
such as AWS Bedrock and GCP Vertex.
Moving response translation into the upstream AI Gateway ext-proc could affect the current retry, fallback, streaming, metrics, and token-accounting lifecycle. The preferred solution is therefore to preserve the existing translation behavior and allow the caching extension to be placed before the downstream AI Gateway ext-proc in request order.
The desired filter order is: Request:
Client
→ caching ext-proc
→ AI Gateway downstream ext-proc
→ Envoy router
→ AI Gateway upstream ext-proc
→ LLM provider
The reverse response order would then be:
Response:
LLM provider
→ AI Gateway upstream ext-proc, response processing skipped
→ Envoy router
→ AI Gateway downstream ext-proc translates the response
→ caching ext-proc stores the canonical OpenAI response
→ Client
This ordering allows the caching extension to hash the original canonical OpenAI request during request processing and observe the translated canonical OpenAI response during response processing. A cache hit can then return the stored response directly without requiring provider-specific translation.
The proposed AIGatewayExtensionPolicy appears to provide the required mechanism. It inserts a composite-wrapped user ext-proc before the downstream AI Gateway ext-proc and enables it on the relevant generated routes, including the initial catch-all route used before x-ai-eg-model is derived. Since Clear Route Cache does not restart filters that have already executed, the caching extension should run once during request processing and receive the response later through the reverse filter path.
To support response caching, the policy must allow response processing rather than using response: Skip. Streaming LLM responses should use streamed response- body processing so that the extension can forward translated SSE chunks while accumulating them for the eventual cache write.
[optional Relevant Links:]
https://github.com/siddharth1036/ai-gateway/blob/91bfb1679cd71fa2b17e0954b3e606becb8c44c8/docs/proposals/012-aigatewayroutefilter/proposal.md
#2364
#2671