Skip to content

fix: let FailOpen InferencePools fall back to another endpoint - #2768

Open
YeiSimon wants to merge 3 commits into
theagentrouter:mainfrom
YeiSimon:fix/inferencepool-failopen
Open

YeiSimon wants to merge 3 commits into
theagentrouter:mainfrom
YeiSimon:fix/inferencepool-failopen

Conversation

@YeiSimon

@YeiSimon YeiSimon commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Description

Make InferencePool failureMode: FailOpen actually fail open, and let a retry leave a failed endpoint.

Today every InferencePool cluster is ORIGINAL_DST on x-gateway-destination-endpoint, and the EPP ext_proc filter has failure_mode_allow: false, so FailOpen behaves like FailClose. Setting only failure_mode_allow (the approach of the closed #2517) is not enough:

  • with the EPP down, the cluster has no destination (503), and a client-set x-gateway-destination-endpoint then picks the upstream;
  • a retry (previous_hosts) cannot leave the endpoint the EPP picked, since the cluster has only that one host;
  • with the default FULL_DUPLEX_STREAMED body mode, Envoy's ext_proc cannot fail open once it has received the request body, and a connection failure to the EPP is only reported after that, so failure_mode_allow has no effect and requests still fail with a 500.

For FailOpen pools:

  • The InferencePool controller creates and owns a headless fallback Service (<pool>-epp-fallback) with the pool's selector and target ports. It is recreated if deleted, deleted when the pool switches to FailClose, and a same-named Service the pool does not control is never modified.
  • The pool's cluster is STRICT_DNS on that Service with an override_host LB policy whose source is the EPP's envoy.lb dynamic metadata only (never a client header), with a LeastRequest fallback for retries and for requests the EPP did not answer. The Service is re-resolved every second.
  • The cluster ejects an endpoint after 3 consecutive connection failures (outlier detection). Application 5xx, 502/503/504 and success rate do not count, so an overloaded but healthy Pod (for example vLLM answering 503) is never ejected, and at most half of the endpoints are ejected.
  • The EPP ext_proc filter sets failure_mode_allow: true and receives the envoy.lb metadata namespace.
  • The EPP's ext_proc cluster gets a TCP active health check (1s interval, also when idle) and panic mode disabled. A dead EPP then leaves the cluster with no healthy host, the ext_proc stream fails in decodeHeaders before any body is received, and failure_mode_allow applies in the default duplex mode.
  • failureMode travels as an optional 7th route/cluster metadata field, written only for FailOpen. FailClose pools keep the 6-field format and the parser reads 6 fields (or an unknown 7th value) as FailClose, so mixed versions during a rolling upgrade are safe.

FailClose pools and pools with failureMode unset are unchanged.

Related Issues/PRs (if applicable)

Fixes #2502
Fixes #2757
Supersedes #2517

Special notes for reviewers (if applicable)

Tested on a kind cluster (Envoy Gateway in gateway namespace mode, Envoy 1.40), default duplex body mode, with the EPP scaled down:

  • requests return 200, ext_proc.failure_mode_allowed increases, and a client-set x-gateway-destination-endpoint pointing outside the pool is ignored;
  • with the EPP up, requests land on the endpoint the EPP picked (with one and with two target ports);
  • a pool Pod that is Ready but refuses connections, with a retry policy: 0 client failures out of 150 requests with the EPP up and with it down (FailClose, as a control: 16, 59 and 30 failures);
  • the pool and its Pods in a different namespace than the Gateway, and an h2c upstream, work with the EPP up and down;
  • a force-killed EPP gave 1 failed request out of 120, a graceful scale-down 0 out of 60; on a freshly started Envoy that has served no request yet, the default no_traffic_interval of 60s would delay detection by up to a minute, which is why the health check sets it to 1s (then 0 to 2 failures);
  • the duplex-mode behavior was also reproduced with local Envoy 1.38.3 and 1.39.1.

The whole InferencePool e2e package passed in 3 consecutive full runs. An earlier full run, right after rebuilding the images, failed once in TestInferencePoolFailOpenTopology (the route did not serve requests for 3 minutes in both subtests); I could not determine the cause, and the same test passed in the 8 other runs I made (alone, in the same order, and in full runs).

New e2e tests cover FailOpen with the EPP up and down, the retry case, the fallback Service lifecycle (recreated if deleted, ports follow targetPorts, a foreign Service is untouched) and the cross-namespace and multi-port cases.

Known limitations and trade-offs:

  • The FailOpen data path depends on the fallback Service. If it is deleted while the controller is not running, the pool has no endpoints (503) until the controller recreates it, even with the EPP up.
  • Switching a pool between FailClose and FailOpen under load can fail a few requests (about 0.1 to 0.3%, within 0.5s of the switch). It needs both the filter and the cluster type to change in one xDS push; changing only one of them gave no errors in 32 switches. I could not determine whether it is inherent to the xDS push or introduced by this change.
  • A request that arrives in the first health check interval after the EPP dies can still fail. A health check blip on a live EPP briefly drops to the fallback load balancing. The health check is TCP on the EPP's gRPC port: it checks reachability and the TLS handshake, not that the EPP is serving.
  • Outlier detection helps only while the EPP is down. While the EPP is up, override_host still uses the EPP's choice even if that endpoint was ejected.
  • The duplex fix relies on Envoy failing the ext_proc stream synchronously when the cluster has no healthy host. The alternative that does not (a passthrough ext_proc as a priority-1 fallback) also worked, but adds a sidecar listener and changes the fallback request to chunked, so I did not take it.
  • Fallback load balancing is LeastRequest, so prefix and KV-cache affinity is lost while the EPP is down. A consistent-hash fallback works with an Envoy Gateway consistentHash policy but needs a hash key header the extproc does not set today; I left it out of this change.
  • Running more than one EPP replica (with a PodDisruptionBudget) reduces how often the fallback is used, but requests in flight on the EPP connection that goes away can still fail.
  • Not tested: Envoy Gateway in its default deployment mode (Envoy in envoy-gateway-system) and a FailOpen pool with a long (hashed) Service name end to end.
  • The topology e2e compares each request with the EPP's own "Request handled" log line, so it depends on the EPP image's log format.

AI Usage

Drafted with Claude Code. I reviewed the change, ran the unit tests, make precommit and the e2e tests locally, and own it.

Generated with Claude Code (https://claude.com/claude-code)

The extension server builds every InferencePool cluster as ORIGINAL_DST on
x-gateway-destination-endpoint and injects the EPP ext_proc filter with
FailureModeAllow false. For a pool with failureMode FailOpen, a retry
therefore always returns to the endpoint that failed, and every request
fails while the endpoint picker is unreachable.

For FailOpen pools:

- The InferencePool controller creates and owns a headless fallback
  Service with the pool's selector and target ports, recreates it if it is
  deleted, and deletes it when the pool switches to FailClose. A Service of
  the same name that the pool doesn't control is never modified, and the
  pool is not accepted while it exists.
- The pool's cluster is STRICT_DNS on that Service with an override_host
  policy. The override comes from the picker's envoy.lb dynamic metadata
  only, never from a client header, with a LeastRequest fallback for
  retries that leave a failed endpoint and for requests the picker didn't
  answer.
- The EPP filter sets failure_mode_allow and receives the envoy.lb
  metadata namespace.
- failureMode travels as a 7th route/cluster metadata field, written only
  for FailOpen pools. FailClose pools keep the 6-field format, and the
  parser reads 6 fields as FailClose, so FailClose pools are unaffected by
  mixed versions during a rolling upgrade.

FailClose pools, and pools with failureMode unset, are unchanged.

Refs theagentrouter#2757, theagentrouter#2502

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YD84biV4jj7buxNzmX4xo2
Signed-off-by: YeiSimon <yeisimon657@gmail.com>
@YeiSimon
YeiSimon requested a review from a team as a code owner October 1, 2026 06:34
@netlify

netlify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for theagentrouter canceled.

Name Link
🔨 Latest commit fbb2dc1
🔍 Latest deploy log https://app.netlify.com/projects/theagentrouter/deploys/6ac1142eb0ca1000089a12c9

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.83673% with 12 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/controller/inference_pool.go 86.53% 7 Missing ⚠️
internal/extensionserver/post_cluster_modify.go 93.15% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

@YeiSimon
YeiSimon marked this pull request as draft October 1, 2026 06:41
YeiSimon and others added 2 commits October 3, 2026 22:41
Envoy's ext_proc cannot fail open in FULL_DUPLEX_STREAMED mode once it has
received the request body, and a connection failure to the endpoint picker is
only reported after that, so failure_mode_allow had no effect with the default
body mode and requests still failed with a 500 when the picker was down.

For FailOpen pools:

- Give the picker's ext_proc cluster a TCP active health check (1s interval,
  also when idle) and disable panic mode. A dead picker then leaves the
  cluster with no healthy host, the ext_proc stream fails in decodeHeaders
  before any body is received, and the request continues to the pool's
  fallback endpoints.
- Eject a fallback endpoint after 3 consecutive connection failures. Application
  5xx and 502/503/504 do not count, so an overloaded Pod (for example vLLM
  answering 503) is not ejected, and at most half of the endpoints are ejected.
- Re-resolve the fallback Service every second, also after a failed lookup.
  When a pool is switched to FailOpen the cluster and the Service are created
  concurrently, and with the default 5s refresh a cluster that resolved the name
  first had no endpoints for up to 5 seconds.

FailClose pools' clusters are unchanged. Metadata with an unknown failureMode
value is read as FailClose.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011ULZDcwVXwrEzHZwWssLJ3
Signed-off-by: YeiSimon <yeisimon657@gmail.com>
Add e2e tests for a FailOpen pool: requests are served with the endpoint
picker up and down, a client-set x-gateway-destination-endpoint is ignored, a
retry leaves an endpoint that is Ready but refuses connections, the fallback
Service is recreated when deleted and follows targetPorts and a Service of the
same name that the pool does not control is not modified, and a pool in another
namespace and a pool with several target ports are served.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011ULZDcwVXwrEzHZwWssLJ3
Signed-off-by: YeiSimon <yeisimon657@gmail.com>
@YeiSimon
YeiSimon marked this pull request as ready for review October 5, 2026 03:30
@missBerg missBerg added bug Something isn't working area/routing Routing, fallback, resilience, InferencePool integration labels Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/routing Routing, fallback, resilience, InferencePool integration bug Something isn't working

Projects

None yet

2 participants