Repository navigation
feat: support for GCP Context Cache - #2784
Draft
sivanantha321 wants to merge 14 commits into
Draft
sivanantha321 wants to merge 14 commits into
sivanantha321 wants to merge 14 commits into
Conversation
Add support for the GCP Vertex AI context caching explicit passthrough path: - Add CachedContent field to gcp.GenerateContentRequest wire schema so the Gemini API receives the cache reference. - Add CachedContent field to GCPVertexAIVendorFields (inline on ChatCompletionRequest) so callers can supply a pre-created cache resource name via the OpenAI-compatible API. - Copy the vendor field into the Gemini request in applyVendorSpecificFields. - Validate that cache_control markers and explicit cachedContent are not used together; return ErrMalformedRequest (HTTP 400) when both are present. - Add gcpRequestHasCacheControlMarkers helper to detect Anthropic-style cache_control on system/user/tool message content parts. - Add unit tests covering passthrough, validation error, and marker detection across all message types. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Introduce the internal/gcpcache package and the GCPAuthHandler interface required to surface GCP credentials to the resolver without duplicating credential management. Changes: - internal/filterapi/runtime.go: add GCPAuthHandler interface extending BackendAuthHandler with GCPTokenSource(), GCPRegion(), GCPProject() methods so in-process components can authenticate to GCP APIs. - internal/backendauth/gcp.go: implement GCPAuthHandler on gcpHandler; add GCPTokenSource() (wraps static token in StaticTokenSource), GCPRegion() and GCPProject() accessors. - internal/translator/gemini_helper.go: export OpenAIMessagesToGeminiContents and OpenAIToolsToGeminiTools so gcpcache can build the cached prefix in Gemini wire format without duplicating conversion logic. - internal/gcpcache/gcpcache.go: new CacheResolver interface and resolver implementation with breakpoint detection, prefix split, deterministic SHA-256 cache key (stored as Google cachedContents displayName), in-memory TTL memo, Google cachedContents list/create client, and post-create re-list to converge duplicate-create races across replicas. - internal/gcpcache/gcpcache_test.go: unit tests covering findBreakpoint, extractTTL, computeCacheKey determinism, cache miss+create, cache hit from Google list, memo hit (no API calls), memo expiry/refetch, create failure, TTL default/override, and duplicate-create race convergence. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Add the GCPCacheSetter interface to the translator package so the GCP Vertex AI translator can receive a resolved cache entry before RequestBody is called. The translator injects the cachedContent resource name into the Gemini request, uses the filtered (non-cached) message remainder as the request contents, and records cache-write token counts in ResponseBody for both streaming and non-streaming responses. Populate a shared CacheResolver on RuntimeBackend (via Server.LoadConfig) for every GCP Vertex AI backend so the in-memory TTL memo is preserved across requests. In upstreamProcessor.ProcessRequestHeaders, after the translator is selected, attempt cache resolution when the resolver, a GCPAuthHandler, and a GCPCacheSetter-capable translator are all present. Resolver failures produce a 502 ImmediateResponse; nil results (no cache_control markers) are a no-op. Add unit tests Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Add a GCPContextCachingSpec to AIServiceBackend in both api/v1beta1 and api/v1alpha1. The field carries Enabled (bool) and DefaultTTL (a GCP seconds-format duration string validated by a kubebuilder regex pattern). The controller wires it into filterapi.Backend.GCPContextCaching so the extproc layer can gate cache-resolver creation on the feature flag. In extproc/server.go, the shared CacheResolver is now only attached when GCPContextCaching.Enabled is true, keeping the resolver nil (and all cachedContents API calls suppressed) for backends that do not opt in. Generated: run make apigen codegen apidoc to regenerate CRD YAML, deepcopy functions, typed clients/listers/informers, and API reference. Tests added: - tests/crdcel: two AIServiceBackend fixtures (valid and invalid defaultTTL) exercised against the live envtest apiserver to verify CRD validation. - internal/controller/gateway_test.go: two reconcileFilterConfigSecret tests asserting that GCPContextCaching is populated when set and nil when absent. - internal/extproc/server_test.go: LoadConfig test asserting resolver creation is gated on Enabled=true / absent field / non-GCP backend. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Update the context cache documentation URL to point to the new docs.cloud.google.com location under gemini-enterprise-agent-platform. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Replace the local createCacheRequest/createCacheResponse/cacheUsageMetadata types in gcpcache with shared types defined in internal/apischema/gcp: - Add CreateCachedContent, CachedContent, CachedContentUsageMetadata, and EncryptionSpec to the gcp schema package so they can be reused across the codebase. - Set CreateCachedContent.TTL as string (GCP duration format e.g. "300s") rather than time.Duration, which marshals as a nanosecond integer and would be rejected by the Vertex AI REST API. - Set CachedContent.ExpireTime as time.Time so the response can be used directly without a separate time.Parse call. - Cast CachedContentUsageMetadata.TotalTokenCount (int32) to int when populating ResolveResult.TokenCount. - Update tests to unmarshal the POST body into gcp.CreateCachedContent instead of the now-removed local type. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The context cache previously kept its resolved cachedContents names in a per-replica in-memory map, so two gateway replicas seeing the same prefix each created their own Google cache. Replace the map with a CacheStore interface and a Redis-backed implementation that all replicas share. This changes the AIServiceBackend API in both v1alpha1 and v1beta1: gcpContextCaching is renamed to contextCache, its enabled field is removed, and a required url field gives the Redis address, either as host:port or as a redis:// URL. Setting contextCache now enables caching; leaving it unset disables it and cache_control markers are ignored. The filterapi config and controller wiring follow the same shape. Creation is arbitrated with SET NX PX: the replica that claims a key creates the cache while the others poll for its result, so a cold prefix arriving at N replicas produces one create rather than N. A waiter that times out resolves against Google itself, which degrades a stuck leader into an extra list rather than a duplicate create. Store failures are never fatal to a request. An unreachable Redis, or a url that does not parse, is treated as a cache miss and the request is served uncached, since caching is an optimization. Google-side failures keep failing fast, because those are actionable by the user: a create rejected for being below the model's minimum token count means the cache_control markers are misplaced, and silently serving the request uncached would hide a cost increase. Resolvers are now carried across config reloads, keyed by backend name and Redis URL, so a filterapi update does not tear down and rebuild each backend's connection pool. Adds github.com/redis/go-redis/v9, and github.com/alicebob/miniredis/v2 for tests. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Creation was arbitrated with a Redis SET NX PX lock: one replica created the cache while the others polled for up to 5s inside ProcessRequestHeaders. The lock was not a guarantee, since a waiter that timed out went on to list and create anyway. It cost a blocking wait on the request path, up to ~100 Redis round trips per waiting request, and timeout constants that were never measured. Resolution on a store miss is now: create the cache, publish it to the store, and return it. There is no waiting on other replicas and no cachedContents.list call. Replicas that miss the same cold prefix at the same moment each create a cache and each keep their own. This is an accepted cost. Once any replica publishes a name, all replicas reuse it. Every create now reports Created with its real TokenCount. Previously the post-create re-list could switch to another replica's cache and report Created=false, which dropped a billed write from the cache-write token metrics. Duplicate creates now show up there. A create response without an expireTime is now an error. Its store TTL would be negative, so the write was skipped silently and every later request for that prefix created again. A leftover lock sentinel from an older replica fails to decode and is treated as a miss, so rolling deploys need no migration. The list code is removed from the request path. A background reconciler that repairs drift between Redis and GCP will follow separately. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The request path no longer lists cachedContents, so a Google cache that the store has forgotten (after a Redis flush or eviction, a dropped write, or a cache created before Redis was configured) is not found again. Add a reconciler that repairs this off the request path. Each round lists cachedContents for one project and region and writes every gateway-created entry to the store with SET NX and a TTL pinned to its expireTime. SET NX means it only fills gaps and never overwrites a name that a request just published. Entries that are not keyed by a cache key were created by some other client and are skipped, as are entries near or past expiry and entries whose expireTime does not parse. The list API has no server-side filter, so the walk reads every entry. It requests the maximum page size, follows nextPageToken, treats a repeated token as the end, and stops after a fixed number of pages. A per-project/region gate, taken with SET NX and the interval as TTL, limits listing to one replica per interval across the fleet without leader election. A missed round has no effect on correctness. The reconciler never serves requests and never returns an error; a failed round is logged and the next one retries. Stores that cannot back it, such as the no-op store, have no reconciler. Nothing starts the reconciler yet; that comes with the resolver lifecycle change. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The shared store and the reconcile loop do not depend on any provider API, but they lived in gcpcache. Move them into a provider-neutral internal/contextcache package so a second provider only has to supply its own API calls. contextcache now holds Store, Entry, NoopStore, the Redis store, and a ReconcileStore interface with SetNX and AcquireGate for stores that can back a reconciler. It also holds Reconciler, which runs the ticker, the per-scope gate, and the SET NX writes, and StaleThreshold. A provider plugs in through Source: GateKey names the scope's gate, and List reports every gateway-created cache in the scope with its cache key. gcpcache keeps the resolver, the cache key, createCache, and a Source that walks cachedContents.list with pagination and the page limit. That source now drops caches created by other clients, and caches whose expireTime does not parse, before the reconciler sees them, so they no longer count as seen or skipped in round stats. Redis keys, the stored value format, and the gcpcache:reconcile: gate prefix are unchanged, so old and new replicas interoperate during a rolling deploy. Store errors now start with "contextcache:" instead of "gcpcache:". Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The reconciler existed but nothing ran it. Tie its lifetime to the resolver that owns it. The resolver gains Start and Close behind a new Background interface, so CacheResolver and the request path are unchanged. Start builds the reconciler for the backend's project and region and runs it under a context the resolver owns. It cannot use the LoadConfig context: the config watcher cancels that when each reload finishes, which would stop the reconciler within seconds. Start does nothing if the resolver is already running, has been closed, or has a store that cannot reconcile. Close cancels the context and waits for the goroutine to exit, and is safe to call more than once. LoadConfig calls Start with the backend's GCPAuthHandler. Resolvers carried across a reload are already running and are not restarted. After a reload, LoadConfig closes every resolver that was not carried over, because its backend was removed, lost its cache config, or was repointed at another Redis. Without this, each such reload would leak a goroutine. Server.Close closes all resolvers, and extproc calls it on shutdown after the gRPC server stops. Rounds run every 60s, starting as soon as the reconciler starts. A round against an unreachable Redis fails and is logged at warn level. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
… path Package layout: - Move gcpcache to internal/contextcache/gcpcache and the Redis store to internal/contextcache/redis (redis.NewStore). contextcache itself now holds only interfaces and shared types, so filterapi can import it without linking go-redis. - Move CacheResolver and ResolveResult into contextcache. Resolve no longer takes an auth handler; the GCP resolver uses the credentials passed to New. - Move CacheSyncer into contextcache with Start() taking no arguments, and fold the separate reconcile-store interface into Store. NoopStore implements the sync methods as no-ops. - Rename the reconciler to syncer and its gate prefix to gcpcache:sync:. - Type filterapi.RuntimeBackend.CacheResolver as contextcache.CacheResolver instead of any. Behavior: - Remove the sync walk's fixed page cap; it now follows nextPageToken until it is empty. - Refresh credentials on reused resolvers: LoadConfig calls SetAuth on every reload because the controller rotates the access token. Resolve and the syncer read the current credentials under a lock. - Include GCP project and region in the cache key, so backends in different locations sharing one Redis do not reuse each other's caches. Existing keys change, so each prefix misses once after deploy. - Skip resolution when a request names its own cachedContent, so no billed cache is created for a request the translator rejects. - Only bill a cache write when the cache was created. - Send expireTime on create only when set: it is now a pointer, since omitempty does not omit a zero time.Time, which was being sent alongside ttl. Naming: - Rename translator GCPCacheSetter/GCPCacheResult to ContextCacheSetter/ContextCacheResult, and pendingCacheResult to cacheResult. Docs and tests: - Drop the claim that the shared store prevents duplicate creates from the ContextCacheSpec.URL and filterapi godoc, and regenerate the CRD and API docs. - Fix the crdcel fixtures, which still used the removed gcpContextCaching field and so did not exercise context caching, and add a missing-url case. - Fix a resolver test that passed only because it had no credentials. Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
✅ Deploy Preview for theagentrouter ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
yuzisun
reviewed
Oct 5, 2026
| // | ||
| // +optional | ||
| // +kubebuilder:validation:Pattern=`^[1-9][0-9]*s$` | ||
| DefaultTTL string `json:"defaultTTL,omitempty"` |
Contributor
There was a problem hiding this comment.
this can be the configuration on the extproc instead as it is more model specific.
yuzisun
reviewed
Oct 5, 2026
| // sent directly to Gemini without calling any cache resolution logic. | ||
| // | ||
| // https://cloud.google.com/vertex-ai/docs/context-cache/context-cache-overview | ||
| CachedContent string `json:"cachedContent,omitzero"` |
Contributor
There was a problem hiding this comment.
user does not have to set this field, it is set by the translator to the request sent to GCP
yuzisun
reviewed
Oct 5, 2026
| SystemInstruction: systemInstruction, | ||
| Tools: tools, | ||
| } | ||
| b, err := json.Marshal(input) |
Contributor
There was a problem hiding this comment.
Does this marshall sort the keys ?
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This commit adds context caching for GCP Vertex AI (Gemini) backends. Clients mark a
prefix of the conversation with Anthropic-style
cache_controlmarkers, as they alreadycan for Claude. The gateway creates a Vertex AI
cachedContentsentry for that prefix,reuses it on later requests, and sends Gemini only the messages after the prefix.
Unlike Anthropic, Vertex AI does not cache implicitly from markers: a cache is a separate
resource that must be created, named, and referenced. The gateway does that work and
records which cache each prefix maps to in a shared Redis, so every replica can reuse a
cache that any replica created.
Configuration
A new optional
contextCachefield onAIServiceBackend, in bothv1alpha1andv1beta1:Setting
contextCacheturns caching on; leaving it unset turns it off. Requests withoutcache_controlmarkers are unaffected. A request can also name its own cache through thecachedContentvendor field; combining that withcache_controlmarkers is rejected.Architecture
flowchart LR C[Client] -->|OpenAI request with cache_control| E[Envoy] E <-->|ext_proc| X subgraph X["extproc"] P[Upstream processor] -->|Resolve| R[gcpcache.Resolver] P -->|SetContextCacheResult| T[Gemini translator] S[Syncer goroutine<br/>one per resolver] end R <-->|GET / SET| RD[(Redis)] S <-->|SET NX| RD R -->|cachedContents.create| G[Vertex AI] S -->|cachedContents.list| G T -->|generateContent<br/>with cachedContent| GRequest path
sequenceDiagram participant P as Upstream processor participant R as Resolver participant RD as Redis participant G as Vertex AI participant T as Translator P->>R: Resolve(request) R->>R: find last cache_control breakpoint,<br/>split prefix / remainder,<br/>key = SHA-256(project, region, model, prefix, tools) R->>RD: GET key alt hit and not within 10s of expiry RD-->>R: cache name R-->>P: name, remainder, Created=false else miss, stale, or Redis error R->>G: cachedContents.create (displayName = key, ttl) G-->>R: name, expireTime, token count R->>RD: SET key, TTL = until expireTime R-->>P: name, remainder, Created=true, TokenCount end P->>T: SetContextCacheResult T->>G: generateContent(remainder, cachedContent = name)The request path never waits on another replica and never lists caches. A cache write is
reported as
cache_creation_input_tokensonly on the request that created it.Background syncer
Redis can lose entries for caches that still exist in GCP: a flush, an eviction, or a
dropped write. Each resolver runs a syncer that repairs this off the request path. Every
60 seconds it:
SET NX gcpcache:sync:<project>/<region>, TTL 60s, soone replica lists each project and region per interval. A replica that loses the gate
skips the round.
cachedContents.list(pageSize=1000, followingnextPageToken). The API hasno server-side filter, so every cache in the project and region is read.
SET NX, so it only fills gaps and neveroverwrites a name a request just published.
The syncer never serves requests. If it fails, requests still work and simply create
more caches.
Lifecycle across config reloads
sequenceDiagram participant W as Config watcher participant S as Server.LoadConfig participant Old as Resolver (removed backend) participant N as Resolver (new or reused) W->>S: reload S->>S: rebuild auth handlers alt backend name + Redis URL already known S->>N: SetAuth(new handler) else new S->>N: gcpcache.New(store, handler) end S->>N: Start() (no-op if running) S->>Old: Close() → syncer exitsResolvers are keyed by backend name and Redis URL and reused across reloads, so a reload
does not rebuild Redis connection pools. A reused resolver receives the new auth handler,
because the controller rotates the GCP access token. A resolver whose backend was removed,
lost its cache config, or moved to another Redis is closed.
Server.Close()stops allsyncers on shutdown.
Failure behavior
contextCache.urldoes not parsecachedContents.createfails (e.g. prefix below the model's minimum token count)expireTimecachedContentandcache_controlmarkers in the same requestRedis failures fail open because caching is an optimization. GCP failures fail fast
because they usually mean the markers are misplaced, and serving the request uncached
would hide a cost increase.
Code layout
internal/contextcacheStore,Entry,NoopStore,CacheResolver,ResolveResult,CacheSyncer,Source, and theSyncerloopinternal/contextcache/redisStore(redis.NewStore)internal/contextcache/gcpcacheResolver: breakpoint and TTL handling, cache key,cachedContents.create, and thecachedContents.listsource for the syncerinternal/translatorContextCacheSetterinterface, implemented by the Gemini translatorinternal/extprocLoadConfig, resolution in the upstream processorNew dependencies:
github.com/redis/go-redis/v9, andgithub.com/alicebob/miniredis/v2for tests.
Related Issues/PRs (if applicable)
Special notes for reviewers (if applicable)
the same moment each create a cache. Each keeps the one it created, and each create is
billed