feat: add configurable request and response guardrails - #2696
fernandoescolar wants to merge 25 commits into
Conversation
✅ Deploy Preview for theagentrouter ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
I haven't checked, as there are many PRs open at the moment, but have you submitted a design doc PR for this? Would be helpful! https://builtonenvoy.io/extensions/bedrock-guardrails/ |
825835e to
474d823
Compare
ee3754a to
6f33ff0
Compare
- add GuardrailPolicy CRD in v1alpha1 and v1beta1 - support dual-version schema and runtime config wiring - compile regex-based guardrails into runtime matcher state - enforce request/response guardrails in extproc before upstream or client response - add guardrail regression tests for runtime evaluation Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
- Implemented support for regex guardrails and external providers including Presidio, AWS Bedrock, and Azure AI Content Safety. - Added failure behavior configuration for external providers. - Enhanced observability with structured logging for guardrail evaluations. - Created test cases for various guardrail policy scenarios including validation of rules and provider configurations. - Added documentation for content guardrails and examples for usage. Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
…ed documentation Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
- Introduced maxRequestBodyBytes and maxResponseBodyBytes to limit payload sizes. - Updated GuardrailEvaluator interface to return GuardrailEvaluationResult. - Enhanced Presidio and Bedrock evaluators to support masking functionality. Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
eea5b29 to
837affe
Compare
|
I’ve added a proposal/technical design document ( I’m sorry I did not follow the project’s contribution and review protocol correctly from the beginning. I’ve reviewed the guidelines and updated the PR to align with them. |
Signed-off-by: Fernando Escolar <fer.escolar@gmail.com>
|
Would love to see native guardrail support! Any thoughts on extensibility? E.g. if users want to add another (not yet supported) guardrail service with a HTTP interface, how could they do that? (I see that you mention calling external executables as being out of scope, but also assume extensibility should have some role in this proprosal). |
|
Yes, I think extensibility should definitely be part of the proposal. One possible approach would be to define a small HTTP interface that custom guardrail services can implement, similar in spirit to Presidio’s API. For example, the runtime could send something like: {
"text": "some user input",
"context": {
"stage": "input"
}
}to a configured endpoint such as: POST /analyzeand expect a normalized response: {
"action": "allow",
"findings": [
{
"type": "PII",
"start": 10,
"end": 20,
"score": 0.92
}
]
}action could be something like allow, block, or modify, with optional findings / replacement content depending on the guardrail. Presidio is a useful precedent here because the analyzer is exposed behind a simple HTTP contract, while the implementation of the detection remains completely independent. That would let us provide native integrations for common guardrail providers, while also having a generic HTTP guardrail adapter for services we don’t support yet. It also avoids the security / portability concerns of invoking arbitrary local executables. |
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Hi! I have added a custom HTTP provider to the PR :) |
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fer.escolar@gmail.com>
Signed-off-by: Fernando Escolar <fernando.escolar@digits.schwarz>
Signed-off-by: Fernando Escolar <fer.escolar@gmail.com>
Description
This change adds native content guardrails to Envoy AI Gateway through a new
GuardrailPolicyresource and request/response enforcement in the external processor.The design proposal is included at
docs/proposals/013-guardrails/proposal.md. It compares the ext-proc architecture with the Built on Envoy Bedrock Guardrails and Azure Content Safety dynamic-module extensions and records the adopted decisions, current scope, follow-up work, and open questions.The API is available in both
v1alpha1andv1beta1and targetsAIServiceBackendresources. Rules support request or response evaluation, deterministic multi-policy composition, configurable payload limits and timeouts, and explicit fail-open or fail-closed behavior. Actions can block, monitor, or mask matching content. CRD admission validates provider configuration, actions, endpoints, targets, and unique rule names.The runtime supports:
Provider credentials are resolved from Kubernetes Secrets or, for Bedrock, from the standard AWS credential chain. Secret changes and policy deletion trigger reconciliation of affected routes. Missing or revoked provider configuration follows the rule's failure mode and does not leave stale gateway configuration active.
External providers receive schema-extracted text instead of serialized request envelopes. Mask actions write transformed text back to its original JSON field. Response guardrails keep streaming responses buffered until evaluation completes, preventing unsafe content from being partially delivered. Violations return HTTP 403 with the
GuardrailViolationerror type.The implementation adds structured logs without payload content, an OpenTelemetry evaluation counter, and guardrail span events. It also includes example manifests and task-oriented documentation.
Validation includes unit tests, CRD admission tests, controller lifecycle and Secret rotation tests, HTTP-stub integration tests, Envoy dataplane tests for request and response blocking, and a Testcontainers test against the pinned official Presidio analyzer image. Live Azure AI Content Safety and AWS Bedrock Guardrails tests passed against real provider resources. Credential-gated live tests remain skipped by default.
GitHub Copilot was used to assist with implementation, tests, documentation, and review. The contributor has reviewed and takes responsibility for the resulting changes.
Related Issues/PRs (if applicable)
fixes: #1409 (Prompt guard moderation)
fixes: #1415 (pii detection)
fixes: #1919 (external security provider)
related: #2128 (content governance)
Special notes for reviewers (if applicable)
The main review areas are the dual-version CRD contract, Secret lifecycle behavior, backend scoping, fail-open/fail-closed semantics, response buffering, and external provider request formats.
Live provider validation requires local credentials and is intentionally not part of default CI. Presidio requires Docker but no credentials.