Skip to content

feat: support for GCP Context Cache - #2784

Draft
sivanantha321 wants to merge 14 commits into
theagentrouter:mainfrom
sivanantha321:gcp-context-cache
Draft

sivanantha321 wants to merge 14 commits into
theagentrouter:mainfrom
sivanantha321:gcp-context-cache

Conversation

@sivanantha321

@sivanantha321 sivanantha321 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Description
This commit adds context caching for GCP Vertex AI (Gemini) backends. Clients mark a
prefix of the conversation with Anthropic-style cache_control markers, as they already
can for Claude. The gateway creates a Vertex AI cachedContents entry for that prefix,
reuses it on later requests, and sends Gemini only the messages after the prefix.

Unlike Anthropic, Vertex AI does not cache implicitly from markers: a cache is a separate
resource that must be created, named, and referenced. The gateway does that work and
records which cache each prefix maps to in a shared Redis, so every replica can reuse a
cache that any replica created.

Configuration

A new optional contextCache field on AIServiceBackend, in both v1alpha1 and
v1beta1:

apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
  name: gemini
spec:
  schema:
    name: GCPVertexAI
  backendRef:
    name: gcp-vertex
    kind: Backend
    group: gateway.envoyproxy.io
  contextCache:
    url: redis.default.svc.cluster.local:6379 # required; host:port or redis:// URL
    defaultTTL: 600s                          # optional; default 300s

Setting contextCache turns caching on; leaving it unset turns it off. Requests without
cache_control markers are unaffected. A request can also name its own cache through the
cachedContent vendor field; combining that with cache_control markers is rejected.

Architecture

flowchart LR
    C[Client] -->|OpenAI request with cache_control| E[Envoy]
    E <-->|ext_proc| X

    subgraph X["extproc"]
        P[Upstream processor] -->|Resolve| R[gcpcache.Resolver]
        P -->|SetContextCacheResult| T[Gemini translator]
        S[Syncer goroutine<br/>one per resolver]
    end

    R <-->|GET / SET| RD[(Redis)]
    S <-->|SET NX| RD
    R -->|cachedContents.create| G[Vertex AI]
    S -->|cachedContents.list| G
    T -->|generateContent<br/>with cachedContent| G
Loading

Request path

sequenceDiagram
    participant P as Upstream processor
    participant R as Resolver
    participant RD as Redis
    participant G as Vertex AI
    participant T as Translator

    P->>R: Resolve(request)
    R->>R: find last cache_control breakpoint,<br/>split prefix / remainder,<br/>key = SHA-256(project, region, model, prefix, tools)
    R->>RD: GET key
    alt hit and not within 10s of expiry
        RD-->>R: cache name
        R-->>P: name, remainder, Created=false
    else miss, stale, or Redis error
        R->>G: cachedContents.create (displayName = key, ttl)
        G-->>R: name, expireTime, token count
        R->>RD: SET key, TTL = until expireTime
        R-->>P: name, remainder, Created=true, TokenCount
    end
    P->>T: SetContextCacheResult
    T->>G: generateContent(remainder, cachedContent = name)
Loading

The request path never waits on another replica and never lists caches. A cache write is
reported as cache_creation_input_tokens only on the request that created it.

Background syncer

Redis can lose entries for caches that still exist in GCP: a flush, an eviction, or a
dropped write. Each resolver runs a syncer that repairs this off the request path. Every
60 seconds it:

  1. Takes a fleet-wide gate with SET NX gcpcache:sync:<project>/<region>, TTL 60s, so
    one replica lists each project and region per interval. A replica that loses the gate
    skips the round.
  2. Walks cachedContents.list (pageSize=1000, following nextPageToken). The API has
    no server-side filter, so every cache in the project and region is read.
  3. Writes each gateway-created entry with SET NX, so it only fills gaps and never
    overwrites a name a request just published.

The syncer never serves requests. If it fails, requests still work and simply create
more caches.

Lifecycle across config reloads

sequenceDiagram
    participant W as Config watcher
    participant S as Server.LoadConfig
    participant Old as Resolver (removed backend)
    participant N as Resolver (new or reused)

    W->>S: reload
    S->>S: rebuild auth handlers
    alt backend name + Redis URL already known
        S->>N: SetAuth(new handler)
    else new
        S->>N: gcpcache.New(store, handler)
    end
    S->>N: Start() (no-op if running)
    S->>Old: Close() → syncer exits
Loading

Resolvers are keyed by backend name and Redis URL and reused across reloads, so a reload
does not rebuild Redis connection pools. A reused resolver receives the new auth handler,
because the controller rotates the GCP access token. A resolver whose backend was removed,
lost its cache config, or moved to another Redis is closed. Server.Close() stops all
syncers on shutdown.

Failure behavior

Condition Result
Redis unreachable, or GET/SET fails Logged; request served uncached (200)
contextCache.url does not parse Logged at reload; backend served uncached
Syncer list or write fails Logged; next round retries
cachedContents.create fails (e.g. prefix below the model's minimum token count) 502
Create response has no expireTime 502, rather than silently never storing the entry
cachedContent and cache_control markers in the same request Rejected as malformed

Redis failures fail open because caching is an optimization. GCP failures fail fast
because they usually mean the markers are misplaced, and serving the request uncached
would hide a cost increase.

Code layout

Package Contents
internal/contextcache Provider-neutral interfaces and types: Store, Entry, NoopStore, CacheResolver, ResolveResult, CacheSyncer, Source, and the Syncer loop
internal/contextcache/redis Redis Store (redis.NewStore)
internal/contextcache/gcpcache GCP Resolver: breakpoint and TTL handling, cache key, cachedContents.create, and the cachedContents.list source for the syncer
internal/translator Optional ContextCacheSetter interface, implemented by the Gemini translator
internal/extproc Resolver lifecycle in LoadConfig, resolution in the upstream processor

New dependencies: github.com/redis/go-redis/v9, and github.com/alicebob/miniredis/v2
for tests.

Related Issues/PRs (if applicable)

Special notes for reviewers (if applicable)

  • Duplicate creates are accepted. Replicas that miss the same new prefix at
    the same moment each create a cache. Each keeps the one it created, and each create is
    billed

Add support for the GCP Vertex AI context caching explicit passthrough path:

- Add CachedContent field to gcp.GenerateContentRequest wire schema so the
  Gemini API receives the cache reference.
- Add CachedContent field to GCPVertexAIVendorFields (inline on
  ChatCompletionRequest) so callers can supply a pre-created cache resource
  name via the OpenAI-compatible API.
- Copy the vendor field into the Gemini request in applyVendorSpecificFields.
- Validate that cache_control markers and explicit cachedContent are not used
  together; return ErrMalformedRequest (HTTP 400) when both are present.
- Add gcpRequestHasCacheControlMarkers helper to detect Anthropic-style
  cache_control on system/user/tool message content parts.
- Add unit tests covering passthrough, validation error, and marker detection
  across all message types.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Introduce the internal/gcpcache package and the GCPAuthHandler interface
required to surface GCP credentials to the resolver without duplicating
credential management.

Changes:
- internal/filterapi/runtime.go: add GCPAuthHandler interface extending
  BackendAuthHandler with GCPTokenSource(), GCPRegion(), GCPProject()
  methods so in-process components can authenticate to GCP APIs.
- internal/backendauth/gcp.go: implement GCPAuthHandler on gcpHandler;
  add GCPTokenSource() (wraps static token in StaticTokenSource),
  GCPRegion() and GCPProject() accessors.
- internal/translator/gemini_helper.go: export OpenAIMessagesToGeminiContents
  and OpenAIToolsToGeminiTools so gcpcache can build the cached prefix in
  Gemini wire format without duplicating conversion logic.
- internal/gcpcache/gcpcache.go: new CacheResolver interface and resolver
  implementation with breakpoint detection, prefix split, deterministic
  SHA-256 cache key (stored as Google cachedContents displayName),
  in-memory TTL memo, Google cachedContents list/create client, and
  post-create re-list to converge duplicate-create races across replicas.
- internal/gcpcache/gcpcache_test.go: unit tests covering findBreakpoint,
  extractTTL, computeCacheKey determinism, cache miss+create, cache hit
  from Google list, memo hit (no API calls), memo expiry/refetch, create
  failure, TTL default/override, and duplicate-create race convergence.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Add the GCPCacheSetter interface to the translator package so the GCP
Vertex AI translator can receive a resolved cache entry before RequestBody
is called. The translator injects the cachedContent resource name into the
Gemini request, uses the filtered (non-cached) message remainder as the
request contents, and records cache-write token counts in ResponseBody for
both streaming and non-streaming responses.

Populate a shared CacheResolver on RuntimeBackend (via Server.LoadConfig)
for every GCP Vertex AI backend so the in-memory TTL memo is preserved
across requests. In upstreamProcessor.ProcessRequestHeaders, after the
translator is selected, attempt cache resolution when the resolver, a
GCPAuthHandler, and a GCPCacheSetter-capable translator are all present.
Resolver failures produce a 502 ImmediateResponse; nil results (no
cache_control markers) are a no-op.

Add unit tests

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Add a GCPContextCachingSpec to AIServiceBackend in both api/v1beta1 and
api/v1alpha1. The field carries Enabled (bool) and DefaultTTL (a GCP
seconds-format duration string validated by a kubebuilder regex pattern).
The controller wires it into filterapi.Backend.GCPContextCaching so the
extproc layer can gate cache-resolver creation on the feature flag.

In extproc/server.go, the shared CacheResolver is now only attached when
GCPContextCaching.Enabled is true, keeping the resolver nil (and all
cachedContents API calls suppressed) for backends that do not opt in.

Generated: run make apigen codegen apidoc to regenerate CRD YAML,
deepcopy functions, typed clients/listers/informers, and API reference.

Tests added:
- tests/crdcel: two AIServiceBackend fixtures (valid and invalid defaultTTL)
  exercised against the live envtest apiserver to verify CRD validation.
- internal/controller/gateway_test.go: two reconcileFilterConfigSecret tests
  asserting that GCPContextCaching is populated when set and nil when absent.
- internal/extproc/server_test.go: LoadConfig test asserting resolver creation
  is gated on Enabled=true / absent field / non-GCP backend.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Update the context cache documentation URL to point to the new
docs.cloud.google.com location under gemini-enterprise-agent-platform.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Replace the local createCacheRequest/createCacheResponse/cacheUsageMetadata
types in gcpcache with shared types defined in internal/apischema/gcp:

- Add CreateCachedContent, CachedContent, CachedContentUsageMetadata, and
  EncryptionSpec to the gcp schema package so they can be reused across
  the codebase.
- Set CreateCachedContent.TTL as string (GCP duration format e.g. "300s")
  rather than time.Duration, which marshals as a nanosecond integer and
  would be rejected by the Vertex AI REST API.
- Set CachedContent.ExpireTime as time.Time so the response can be used
  directly without a separate time.Parse call.
- Cast CachedContentUsageMetadata.TotalTokenCount (int32) to int when
  populating ResolveResult.TokenCount.
- Update tests to unmarshal the POST body into gcp.CreateCachedContent
  instead of the now-removed local type.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The context cache previously kept its resolved cachedContents names in a
per-replica in-memory map, so two gateway replicas seeing the same prefix
each created their own Google cache. Replace the map with a CacheStore
interface and a Redis-backed implementation that all replicas share.

This changes the AIServiceBackend API in both v1alpha1 and v1beta1:
gcpContextCaching is renamed to contextCache, its enabled field is
removed, and a required url field gives the Redis address, either as
host:port or as a redis:// URL. Setting contextCache now enables
caching; leaving it unset disables it and cache_control markers are
ignored. The filterapi config and controller wiring follow the same
shape.

Creation is arbitrated with SET NX PX: the replica that claims a key
creates the cache while the others poll for its result, so a cold prefix
arriving at N replicas produces one create rather than N. A waiter that
times out resolves against Google itself, which degrades a stuck leader
into an extra list rather than a duplicate create.

Store failures are never fatal to a request. An unreachable Redis, or a
url that does not parse, is treated as a cache miss and the request is
served uncached, since caching is an optimization. Google-side failures
keep failing fast, because those are actionable by the user: a create
rejected for being below the model's minimum token count means the
cache_control markers are misplaced, and silently serving the request
uncached would hide a cost increase.

Resolvers are now carried across config reloads, keyed by backend name
and Redis URL, so a filterapi update does not tear down and rebuild each
backend's connection pool.

Adds github.com/redis/go-redis/v9, and github.com/alicebob/miniredis/v2
for tests.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Creation was arbitrated with a Redis SET NX PX lock: one replica created
the cache while the others polled for up to 5s inside
ProcessRequestHeaders. The lock was not a guarantee, since a waiter that
timed out went on to list and create anyway. It cost a blocking wait on
the request path, up to ~100 Redis round trips per waiting request, and
timeout constants that were never measured.

Resolution on a store miss is now: create the cache, publish it to the
store, and return it. There is no waiting on other replicas and no
cachedContents.list call. Replicas that miss the same cold prefix at the
same moment each create a cache and each keep their own. This is an
accepted cost. Once any replica publishes a name, all replicas reuse it.

Every create now reports Created with its real TokenCount. Previously
the post-create re-list could switch to another replica's cache and
report Created=false, which dropped a billed write from the cache-write
token metrics. Duplicate creates now show up there.

A create response without an expireTime is now an error. Its store TTL
would be negative, so the write was skipped silently and every later
request for that prefix created again.

A leftover lock sentinel from an older replica fails to decode and is
treated as a miss, so rolling deploys need no migration.

The list code is removed from the request path. A background reconciler
that repairs drift between Redis and GCP will follow separately.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The request path no longer lists cachedContents, so a Google cache that
the store has forgotten (after a Redis flush or eviction, a dropped
write, or a cache created before Redis was configured) is not found
again. Add a reconciler that repairs this off the request path.

Each round lists cachedContents for one project and region and writes
every gateway-created entry to the store with SET NX and a TTL pinned
to its expireTime. SET NX means it only fills gaps and never overwrites
a name that a request just published. Entries that are not keyed by a
cache key were created by some other client and are skipped, as are
entries near or past expiry and entries whose expireTime does not
parse.

The list API has no server-side filter, so the walk reads every entry.
It requests the maximum page size, follows nextPageToken, treats a
repeated token as the end, and stops after a fixed number of pages.

A per-project/region gate, taken with SET NX and the interval as TTL,
limits listing to one replica per interval across the fleet without
leader election. A missed round has no effect on correctness.

The reconciler never serves requests and never returns an error; a
failed round is logged and the next one retries. Stores that cannot
back it, such as the no-op store, have no reconciler. Nothing starts
the reconciler yet; that comes with the resolver lifecycle change.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The shared store and the reconcile loop do not depend on any provider
API, but they lived in gcpcache. Move them into a provider-neutral
internal/contextcache package so a second provider only has to supply
its own API calls.

contextcache now holds Store, Entry, NoopStore, the Redis store, and a
ReconcileStore interface with SetNX and AcquireGate for stores that can
back a reconciler. It also holds Reconciler, which runs the ticker, the
per-scope gate, and the SET NX writes, and StaleThreshold. A provider
plugs in through Source: GateKey names the scope's gate, and List
reports every gateway-created cache in the scope with its cache key.

gcpcache keeps the resolver, the cache key, createCache, and a Source
that walks cachedContents.list with pagination and the page limit. That
source now drops caches created by other clients, and caches whose
expireTime does not parse, before the reconciler sees them, so they no
longer count as seen or skipped in round stats.

Redis keys, the stored value format, and the gcpcache:reconcile: gate
prefix are unchanged, so old and new replicas interoperate during a
rolling deploy. Store errors now start with "contextcache:" instead of
"gcpcache:".

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
The reconciler existed but nothing ran it. Tie its lifetime to the
resolver that owns it.

The resolver gains Start and Close behind a new Background interface,
so CacheResolver and the request path are unchanged. Start builds the
reconciler for the backend's project and region and runs it under a
context the resolver owns. It cannot use the LoadConfig context: the
config watcher cancels that when each reload finishes, which would stop
the reconciler within seconds. Start does nothing if the resolver is
already running, has been closed, or has a store that cannot reconcile.
Close cancels the context and waits for the goroutine to exit, and is
safe to call more than once.

LoadConfig calls Start with the backend's GCPAuthHandler. Resolvers
carried across a reload are already running and are not restarted.
After a reload, LoadConfig closes every resolver that was not carried
over, because its backend was removed, lost its cache config, or was
repointed at another Redis. Without this, each such reload would leak
a goroutine.

Server.Close closes all resolvers, and extproc calls it on shutdown
after the gRPC server stops.

Rounds run every 60s, starting as soon as the reconciler starts. A
round against an unreachable Redis fails and is logged at warn level.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
… path

Package layout:
- Move gcpcache to internal/contextcache/gcpcache and the Redis store to
  internal/contextcache/redis (redis.NewStore). contextcache itself now
  holds only interfaces and shared types, so filterapi can import it
  without linking go-redis.
- Move CacheResolver and ResolveResult into contextcache. Resolve no
  longer takes an auth handler; the GCP resolver uses the credentials
  passed to New.
- Move CacheSyncer into contextcache with Start() taking no arguments,
  and fold the separate reconcile-store interface into Store. NoopStore
  implements the sync methods as no-ops.
- Rename the reconciler to syncer and its gate prefix to gcpcache:sync:.
- Type filterapi.RuntimeBackend.CacheResolver as
  contextcache.CacheResolver instead of any.

Behavior:
- Remove the sync walk's fixed page cap; it now follows nextPageToken
  until it is empty.
- Refresh credentials on reused resolvers: LoadConfig calls SetAuth on
  every reload because the controller rotates the access token. Resolve
  and the syncer read the current credentials under a lock.
- Include GCP project and region in the cache key, so backends in
  different locations sharing one Redis do not reuse each other's
  caches. Existing keys change, so each prefix misses once after deploy.
- Skip resolution when a request names its own cachedContent, so no
  billed cache is created for a request the translator rejects.
- Only bill a cache write when the cache was created.
- Send expireTime on create only when set: it is now a pointer, since
  omitempty does not omit a zero time.Time, which was being sent
  alongside ttl.

Naming:
- Rename translator GCPCacheSetter/GCPCacheResult to
  ContextCacheSetter/ContextCacheResult, and pendingCacheResult to
  cacheResult.

Docs and tests:
- Drop the claim that the shared store prevents duplicate creates from
  the ContextCacheSpec.URL and filterapi godoc, and regenerate the CRD
  and API docs.
- Fix the crdcel fixtures, which still used the removed
  gcpContextCaching field and so did not exercise context caching, and
  add a missing-url case.
- Fix a resolver test that passed only because it had no credentials.

Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
Signed-off-by: Sivanantham Chinnaiyan <sivanantham.chinnaiyan@ideas2it.com>
@netlify

netlify Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for theagentrouter ready!

Name Link
🔨 Latest commit 34fd948
🔍 Latest deploy log https://app.netlify.com/projects/theagentrouter/deploys/6ac28a022294600008f58bdb
😎 Deploy Preview https://deploy-preview-2784--theagentrouter.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

//
// +optional
// +kubebuilder:validation:Pattern=`^[1-9][0-9]*s$`
DefaultTTL string `json:"defaultTTL,omitempty"`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this can be the configuration on the extproc instead as it is more model specific.

// sent directly to Gemini without calling any cache resolution logic.
//
// https://cloud.google.com/vertex-ai/docs/context-cache/context-cache-overview
CachedContent string `json:"cachedContent,omitzero"`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

user does not have to set this field, it is set by the translator to the request sent to GCP

SystemInstruction: systemInstruction,
Tools: tools,
}
b, err := json.Marshal(input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this marshall sort the keys ?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants