Repository navigation
Conversation
The extension server builds every InferencePool cluster as ORIGINAL_DST on x-gateway-destination-endpoint and injects the EPP ext_proc filter with FailureModeAllow false. For a pool with failureMode FailOpen, a retry therefore always returns to the endpoint that failed, and every request fails while the endpoint picker is unreachable. For FailOpen pools: - The InferencePool controller creates and owns a headless fallback Service with the pool's selector and target ports, recreates it if it is deleted, and deletes it when the pool switches to FailClose. A Service of the same name that the pool doesn't control is never modified, and the pool is not accepted while it exists. - The pool's cluster is STRICT_DNS on that Service with an override_host policy. The override comes from the picker's envoy.lb dynamic metadata only, never from a client header, with a LeastRequest fallback for retries that leave a failed endpoint and for requests the picker didn't answer. - The EPP filter sets failure_mode_allow and receives the envoy.lb metadata namespace. - failureMode travels as a 7th route/cluster metadata field, written only for FailOpen pools. FailClose pools keep the 6-field format, and the parser reads 6 fields as FailClose, so FailClose pools are unaffected by mixed versions during a rolling upgrade. FailClose pools, and pools with failureMode unset, are unchanged. Refs theagentrouter#2757, theagentrouter#2502 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YD84biV4jj7buxNzmX4xo2 Signed-off-by: YeiSimon <yeisimon657@gmail.com>
✅ Deploy Preview for theagentrouter canceled.
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
YeiSimon
marked this pull request as draft
October 1, 2026 06:41
Envoy's ext_proc cannot fail open in FULL_DUPLEX_STREAMED mode once it has received the request body, and a connection failure to the endpoint picker is only reported after that, so failure_mode_allow had no effect with the default body mode and requests still failed with a 500 when the picker was down. For FailOpen pools: - Give the picker's ext_proc cluster a TCP active health check (1s interval, also when idle) and disable panic mode. A dead picker then leaves the cluster with no healthy host, the ext_proc stream fails in decodeHeaders before any body is received, and the request continues to the pool's fallback endpoints. - Eject a fallback endpoint after 3 consecutive connection failures. Application 5xx and 502/503/504 do not count, so an overloaded Pod (for example vLLM answering 503) is not ejected, and at most half of the endpoints are ejected. - Re-resolve the fallback Service every second, also after a failed lookup. When a pool is switched to FailOpen the cluster and the Service are created concurrently, and with the default 5s refresh a cluster that resolved the name first had no endpoints for up to 5 seconds. FailClose pools' clusters are unchanged. Metadata with an unknown failureMode value is read as FailClose. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ULZDcwVXwrEzHZwWssLJ3 Signed-off-by: YeiSimon <yeisimon657@gmail.com>
Add e2e tests for a FailOpen pool: requests are served with the endpoint picker up and down, a client-set x-gateway-destination-endpoint is ignored, a retry leaves an endpoint that is Ready but refuses connections, the fallback Service is recreated when deleted and follows targetPorts and a Service of the same name that the pool does not control is not modified, and a pool in another namespace and a pool with several target ports are served. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ULZDcwVXwrEzHZwWssLJ3 Signed-off-by: YeiSimon <yeisimon657@gmail.com>
YeiSimon
marked this pull request as ready for review
October 5, 2026 03:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Make InferencePool
failureMode: FailOpenactually fail open, and let a retry leave a failed endpoint.Today every InferencePool cluster is
ORIGINAL_DSTonx-gateway-destination-endpoint, and the EPP ext_proc filter hasfailure_mode_allow: false, soFailOpenbehaves likeFailClose. Setting onlyfailure_mode_allow(the approach of the closed #2517) is not enough:x-gateway-destination-endpointthen picks the upstream;previous_hosts) cannot leave the endpoint the EPP picked, since the cluster has only that one host;FULL_DUPLEX_STREAMEDbody mode, Envoy's ext_proc cannot fail open once it has received the request body, and a connection failure to the EPP is only reported after that, sofailure_mode_allowhas no effect and requests still fail with a 500.For
FailOpenpools:<pool>-epp-fallback) with the pool's selector and target ports. It is recreated if deleted, deleted when the pool switches toFailClose, and a same-named Service the pool does not control is never modified.STRICT_DNSon that Service with anoverride_hostLB policy whose source is the EPP'senvoy.lbdynamic metadata only (never a client header), with aLeastRequestfallback for retries and for requests the EPP did not answer. The Service is re-resolved every second.failure_mode_allow: trueand receives theenvoy.lbmetadata namespace.decodeHeadersbefore any body is received, andfailure_mode_allowapplies in the default duplex mode.failureModetravels as an optional 7th route/cluster metadata field, written only forFailOpen.FailClosepools keep the 6-field format and the parser reads 6 fields (or an unknown 7th value) asFailClose, so mixed versions during a rolling upgrade are safe.FailClosepools and pools withfailureModeunset are unchanged.Related Issues/PRs (if applicable)
Fixes #2502
Fixes #2757
Supersedes #2517
Special notes for reviewers (if applicable)
Tested on a kind cluster (Envoy Gateway in gateway namespace mode, Envoy 1.40), default duplex body mode, with the EPP scaled down:
ext_proc.failure_mode_allowedincreases, and a client-setx-gateway-destination-endpointpointing outside the pool is ignored;no_traffic_intervalof 60s would delay detection by up to a minute, which is why the health check sets it to 1s (then 0 to 2 failures);The whole InferencePool e2e package passed in 3 consecutive full runs. An earlier full run, right after rebuilding the images, failed once in
TestInferencePoolFailOpenTopology(the route did not serve requests for 3 minutes in both subtests); I could not determine the cause, and the same test passed in the 8 other runs I made (alone, in the same order, and in full runs).New e2e tests cover FailOpen with the EPP up and down, the retry case, the fallback Service lifecycle (recreated if deleted, ports follow
targetPorts, a foreign Service is untouched) and the cross-namespace and multi-port cases.Known limitations and trade-offs:
FailCloseandFailOpenunder load can fail a few requests (about 0.1 to 0.3%, within 0.5s of the switch). It needs both the filter and the cluster type to change in one xDS push; changing only one of them gave no errors in 32 switches. I could not determine whether it is inherent to the xDS push or introduced by this change.override_hoststill uses the EPP's choice even if that endpoint was ejected.LeastRequest, so prefix and KV-cache affinity is lost while the EPP is down. A consistent-hash fallback works with an Envoy GatewayconsistentHashpolicy but needs a hash key header the extproc does not set today; I left it out of this change.envoy-gateway-system) and a FailOpen pool with a long (hashed) Service name end to end.AI Usage
Drafted with Claude Code. I reviewed the change, ran the unit tests,
make precommitand the e2e tests locally, and own it.Generated with Claude Code (https://claude.com/claude-code)