Test every tool call. Verify every failure.
Replay-verified security evaluation and deterministic policy hardening for AI agents that use Model Context Protocol tools.
Features · How It Works · Quick Start · Documentation
MCP Breaker is a security-evaluation platform for AI agents with MCP tool access. It inventories the tools an agent can call, runs adversarial and safe-control scenarios, captures the resulting tool and state evidence, and replays suspicious behavior from a clean state before promoting it to a verified finding.
Verified findings feed a deterministic least-privilege policy engine. MCP Breaker then reruns the same attacks against the hardened policy and records whether each unsafe action was blocked, approval-gated, or constrained to a sandbox.
Note
The interactive demo runs locally. MCP inspection, TrueForge execution, generated evidence, and policy retests remain intentionally loopback-only because the demo target and control plane bind only to local addresses.
| Capability | What MCP Breaker verifies |
|---|---|
| MCP tool inventory | Discovers tool names, schemas, annotations, and risk-relevant capabilities. |
| Adversarial scenarios | Tests indirect prompt injection, unauthorized writes, destructive actions, and confused-deputy behavior. |
| Safe controls | Confirms that expected read-only behavior still works and distinguishes a secure agent from a broken one. |
| Runtime evidence | Records tool arguments, results, state changes, session identifiers, and expected-versus-observed behavior. |
| Clean-state replay | Reproduces a candidate in a distinct run before it becomes a verified finding. |
| Verified scoring | Calculates security scores only from replay-verified runtime findings. |
| Policy hardening | Generates explicit ALLOW, DENY, REQUIRE_APPROVAL, and SANDBOX_ONLY rules with provenance. |
| Before/after proof | Reruns the suite and records the exact policy decision and remediation result for every finding. |
MCP target
↓
tool discovery + risk classification
↓
adversarial scenarios + safe controls
↓
tool traces + direct state evidence
↓
clean-state replay
↓
verified findings + evidence-backed score
↓
least-privilege policy generation
↓
hardened retest + remediation proof
MCP Breaker supports two evidence paths:
- Deterministic offline mode provides a fast, repeatable regression harness with no model calls.
- TrueForge live mode drives a real model through the disposable MCP target, persists SDK events, reproduces candidates in new sessions, and verifies approval pauses during hardened retests.
The matrix separates verified failures, completed safe executions, candidates, and untested combinations. A tool is never marked safe merely because a scenario did not run.
Every verified finding links the original unsafe trace to a separate replay trace and exposes the exact tool arguments, response, and resulting state.
The hardening view explains each effective rule change and compares the current policy with the live retest policy.
| Layer | Technology |
|---|---|
| Language | TypeScript 6, Node.js 24+ |
| Dashboard | Next.js 16, React 19, Tailwind CSS 4 |
| MCP runtime | @modelcontextprotocol/server, @modelcontextprotocol/client |
| Live evaluation | Official @truefoundry/trueforge-sdk |
| Validation | Zod schemas at artifact and policy boundaries |
| Testing | Vitest, TypeScript project references, ESLint |
| Repository | npm workspaces monorepo |
apps/
├── dashboard/ Next.js security dashboard and local control API
└── demo-target/ Disposable, resettable MCP server with 11 tools
packages/
├── attack-library/ Adversarial and safe-control scenario corpus
├── breaker-core/ Tool traces, policy validation, and enforcement
├── evaluation/ Deterministic and TrueForge execution pipelines
├── scoring/ Verified-finding security scoring
└── shared/ Shared schemas and report contracts
docs/ Architecture, stage notes, and review workflow
tests/ Offline regression and integration coverage
Requirements: Node.js 24 or newer and npm.
git clone https://github.com/Muzzy5150/mcp-breaker.git
cd mcp-breaker
npm install
npm run verify
npm run demoOpen http://127.0.0.1:3000 after the dashboard starts. The deterministic demo does not require model credentials or paid calls.
Configure the local TrueForge control plane, model provider, and Daytona sandbox as described in Stage 5, then run:
npm run demo:uiThe website owns the disposable MCP server on 127.0.0.1:18880. Do not run npm run demo:start at the same time.
# Complete offline verification
npm run verify
# Deterministic assessment and hardening
npm run demo:assessment
npm run demo:hardening
# Live TrueForge assessment and hardening
npm run trueforge:doctor
npm run demo:live
npm run demo:live:hardening
npm run test:liveGenerated assessment artifacts are schema-validated, written under artifacts/, permissioned for local use, and ignored by Git.
- The demo target contains disposable simulated repositories, files, issues, pull requests, and messages.
- MCP and TrueForge endpoints must be loopback-only; arbitrary remote MCP URLs are rejected.
- A failed or missing replay remains a candidate or inconclusive result, never a verified vulnerability.
- Unknown tools fail closed under the hardening policy.
- The harness never kills a process occupying its managed MCP port.
- No credentials, hidden model reasoning, or generated assessment artifacts are committed.
- Architecture and product boundaries
- Deterministic demo target
- Scenario runner and verifier
- Security dashboard
- Policy hardening and retest
- Live TrueForge evaluation
- Website demo flow
- GitHub and Qodo review workflow
Codex was used as a coding assistant during development. Project architecture, verification strategy, tests, security boundaries, and final engineering decisions were reviewed and directed by the participant. The public pull-request history is the primary Qodo review evidence source.



