Skip to content

Repository files navigation

FoFo — Test Integrity for AI Coding Agents

Claude Code plugin test-integrity version 1.0.0-beta.4 MIT license deps stack-agnostic contract

FoFo — Test Integrity for AI Coding Agents

Your AI can't grade its own homework.
A test-integrity layer for AI coding agents: the context that grades the tests is isolated from the context that writes the code, and test quality is enforced by scripts that exit non-zero — not instructions the model is trusted to follow.

Not another spec framework. A verification layer that sits under whatever spec or TDD workflow you already use.

See it pass in ~30s — no API key:
cd plugins/fofo/skills/sdlc/examples/sample-feature && bash smoke.sh → ALL CHECKS PASSED


Why

AI agents write tests that pass. That's the problem. A test the same context wrote to match the code it just wrote proves almost nothing — it locks in behavior without judging it. FoFo fixes the trust of your tests:

  • An independent context (the Judge) writes the tests from the spec, before any code exists.
  • A separate context (the Author) writes code to pass tests it didn't write and cannot silently weaken.
  • A gate runner (the Referee) grades the code and the tests — assertion intent, traceability, separation, secrets — and escalates to a human only what the gates can't resolve.

Procedure lives in the skill. Enforcement lives in scripts with real exit codes, so it can't be summarized away or gamed.

How this differs

If you know TDD-Guard — the closest neighbor — here's the delta. TDD-Guard polices the order of red → green → refactor (no implementation without a failing test, no over-implementation). FoFo adds the two things order alone can't give you:

  1. Context separation. The test-author context and the code-author context are different (separation-gate), so the agent structurally can't write tests with the same blind spots as the code it's about to write. Cycle-ordering can't catch a test authored to match the implementation.
  2. Test-quality grading. intent-gate grades the tests themselves — assertion-free, trivially-true, snapshot-only, and banned-mock tests fail the gate, via a real AST parse — not just "a failing test existed first."

Mutation testing is complementary, not competing. "Tests that pass" ≠ "tests that catch" — mutation testing is the gold standard for the latter. FoFo treats mutation (and coverage, lint) as bring-your-own gates wrapped to the same contract: point it at your mutation tool and it runs in the same pipeline and routes failures the same way. FoFo absorbs that approach rather than replacing it.

Spec-driven development tools (Spec Kit, Kiro, BMAD, AI-DLC) are upstream, not rivals. They own specify → plan → tasks; none of them enforces test integrity. FoFo docks underneath whichever one you use: scripts/spec-adapter converts a Spec Kit spec.md (FR-### bullets), a Kiro requirements.md (EARS criteria), or a BMAD PRD (FR#/NFR# lines) into the REQUIREMENTS.md the loop consumes, and spec-lint accepts EARS-format criteria (WHEN … THEN the system SHALL …) natively. Author the spec wherever you like; verify it here.

Install

In Claude Code:

/plugin marketplace add Sweet-Papa-Technologies/Agentic-SDLC
/plugin install fofo@fofo-marketplace

Then use it with /fofo:sdlc, or just describe a spec-to-code task ("implement this from the spec", "write tests TDD") and Claude loads it automatically. The Judge and Reviewer subagents auto-load — check /agents.

Works on Claude Code out of the box; the same folder is format-compatible with other SKILL.md-consuming agents (Cursor, Codex CLI, Gemini CLI). See INSTALL.md.

How it works

FoFo phase/role flow

Four roles, run as separate contexts where the host allows (see Guarantees & limits):

Role Who Sees Never does
Operator the human everything owns spec sign-off, escalations, security
Judge sub-agent, spec-only spec + test intents see the implementation
Author separate context tests + code edit the tests that grade it
Referee gate scripts the diff make calls scripts can't

When verification fails, the runner tags each escalation with a route: code failures go back to the Author, tests failures go back to the Judge in a fresh pass. The Author never edits the tests that grade it.

The loop

Running /fofo:sdlc choreographs eight phases. Each has an exit gate — nothing proceeds until it passes or the failure is escalated to a human.

Phase Role What happens Exit gate
0 · Spec & Design Operator Each requirement gets a stable ID + acceptance criteria (REQUIREMENTS.md) spec-lint
1 · Test Requirements Judge · spec-only Map every requirement → test intents (TEST-REQS.yaml) trace-gate
2 · Test Authoring Judge · tests-first Write the tests, tagged by requirement ID — red, because no code exists yet — then hash them into TEST-LOCK.json (test-lock --write) redgreen --expect red
3 · Implementation Author · separate ctx Write code to turn the suite green; never touch the tests redgreen --expect green
4 · Verification Referee · Tier 0 Run the gate runner: trace · intent · secret-scan · test-lock · oversight-integrity · deps-gate (+ opt-in mutation/coverage/separation/flake/diff-budget) all hard gates pass, or escalate
5 · Fresh-Eyes Review Reviewer · Tier 1 Model gates: semantic-test-judge + fresh-eyes-review (cheating / security / silent changes) + trajectory-judge (audits the Author's transcript for how the code was produced) clean, or escalate
6 · Human Escalation Operator · Tier 2 Gets only the tight escalation payload — gate, route, findings — not raw diffs Operator decides
7 · Integration — Keep the PR small enough to review (diff-budget), then merge merge

Loop-back (the critical rule): a failed gate routes back to whoever owns the fix — code failures to the Author, tests failures to the Judge in a fresh pass. The Author can never edit the tests that grade it.

What a run produces

The Referee (gate-runner) emits one JSON verdict and one exit code — the shared enforcement surface that makes the loop wireable into CI or a pre-merge hook instead of being advice the model may ignore (abbreviated):

{
  "overall_status": "fail",          // pass · fail · escalate
  "overall_exit": 1,                 // 0 pass · 1 fail · 2 escalate · 3 skip
  "gates": [
    { "gate": "secret-scan", "status": "pass" },
    { "gate": "intent-gate", "status": "fail", "route": "tests",
      "summary": "1 test asserts no intent",
      "findings": [{ "severity": "high", "location": "test/slug.test.js:9",
                     "detail": "Test \"builds a slug\" is assertion-free." }] }
  ],
  "escalation": [
    { "gate": "intent-gate", "route": "tests", "summary": "1 test asserts no intent" }
  ]
}

The gates

Every gate — kept or bring-your-own — obeys one I/O contract (--changed --policy, prints one JSON verdict, exits 0 pass / 1 fail / 2 escalate / 3 skip). Mix and match; point config at your stack.

Gate Tier What it checks Default
spec-lint 0 every requirement has an ID + acceptance criteria ✅ on
trace-gate 0 every requirement maps to ≥1 test; flags orphan tests ✅ on
redgreen-gate 0 suite is red before code, green after ✅ on
intent-gate 0 catches assertion-free / trivially-true / snapshot-only / banned-mock tests — depth varies by language (see below) ✅ on
secret-scan 0 deterministic hard-coded-secret detector (AWS keys, PEM private keys hard-fail; secret-looking assignments escalate) — key-free, no network ✅ on
test-lock 0 tamper evidence: the Judge hashes the tests at Phase 2 exit (--write → TEST-LOCK.json); a locked test modified or deleted afterward hard-fails, a new unlocked test escalates ✅ on
oversight-integrity 0 the diff must not touch the referee's own machinery (gates.config, policy.json, PROVENANCE.*, TEST-LOCK.json) — agents sabotaging their own oversight is a documented failure mode ✅ on
deps-gate 0 a new dependency in any manifest (package.json, requirements.txt, pyproject.toml, go.mod, Cargo.toml, Gemfile) escalates for Operator sign-off — supply-chain surface no test will catch ✅ on
separation-gate 0 fails if one context authored both the tests and the code for a unit ⚙️ opt-in
flake-gate 0 re-runs the suite N times; fails if the result is non-deterministic (a flaky test is an untrustworthy test) ⚙️ opt-in
diff-budget 0 caps changed lines so a PR stays small enough to actually review (Phase 7); git-diff or line-count mode ⚙️ opt-in
eval-gate 0 red/green for nondeterministic (LLM-app) code: runs your eval harness N trials, holds pass rate and optional mean score to a floor — one green run of a probabilistic unit proves nothing ⚙️ opt-in
mutation · coverage · lint 0 bring your own command, wrapped to the contract ⚙️ opt-in
semantic-test-judge 1 model judges whether each test asserts requirement intent vs. just touching lines ⚙️ opt-in
fresh-eyes-review 1 model scans the diff for cheating, security, silent architecture changes ⚙️ opt-in
trajectory-judge 1 model reads the Author's exported transcript for cheat signals (intent to game, test tampering, hardcoded expectations, forbidden channels, misreporting) — judges the process, the strongest counter to reward hacking in 2026 research ⚙️ opt-in

No mutation/coverage/lint/model-proxy/language is hardcoded. Heavy and model gates are disabled by default, so a fresh install runs the language-agnostic gates immediately with nothing else installed.

Guarantees & limits

Two things a careful adopter should know up front. Both are honest seams, not surprises.

Separation is as hard as your host. On a host that can truly isolate contexts — e.g. Claude Code subagents — the Judge literally cannot read the implementation, and separation-gate is a hard guarantee. On a host that can't isolate contexts, separation degrades to a convention, and separation-gate falls back to an after-the-fact heuristic: did the same session author both the tests and the code, per the provenance manifest? Today the hard guarantee holds on Claude Code (subagents); elsewhere it is best-effort.

intent-gate depth varies by language. The contract is stack-agnostic — one I/O shape, point it at any stack — but enforcement depth is not uniform yet:

Language intent-gate depth
JavaScript / TypeScript Full AST (vendored acorn parser) — precise structural detection
JSX / TSX, other languages Token-level heuristics (assertion-free / trivially-true / banned-mock)
Anything else Bring your own parser behind the same JSON contract — see porting

So depth is deepest on JS/TS today and token-level elsewhere until per-language parsers are added. The other Tier-0 gates (spec-lint, trace-gate, redgreen-gate, secret-scan, separation-gate) are genuinely language-agnostic, and redgreen works against any test command you configure.

Quickstart (per repo)

# 1. drop in the config (one time, from the installed plugin dir)
cp "$HOME/.claude/plugins/marketplaces/fofo-marketplace/plugins/fofo/skills/sdlc/gates.config.example" gates.config
cp "$HOME/.claude/plugins/marketplaces/fofo-marketplace/plugins/fofo/skills/sdlc/policy.json" policy.json

# 2. point red/green at your test command (policy.json -> gates.redgreen-gate.test_command)

# 3. run the whole gate runner over your changes
"$HOME/.claude/plugins/.../skills/sdlc/scripts/gate-runner" \
  --config gates.config --changed "src/**/* test/**/*"

Or just run /fofo:sdlc and let the skill choreograph the phases.

What's in the box

Agentic-SDLC/                         # project home AND plugin marketplace
├── .claude-plugin/marketplace.json   # this catalog
├── assets/                           # logo, banner, icons, diagram
└── plugins/fofo/
    ├── .claude-plugin/plugin.json    # plugin manifest
    ├── agents/                       # Judge + Reviewer subagents (auto-load)
    └── skills/sdlc/                  # the skill: SKILL.md, gate scripts, references, runnable smoke tests

Deep docs ship with the skill: SKILL.md · gate contract · phases & roles · where FoFo sits in the ASDLC landscape · porting to other languages · design notes

Verify it yourself

The skill ships runnable end-to-end smoke tests (no API key needed — the model gate is exercised against a local mock):

cd plugins/fofo/skills/sdlc/examples/sample-feature && bash smoke.sh   # the full loop
cd plugins/fofo/skills/sdlc/examples/security-gates  && bash smoke.sh   # secret-scan, dogfooded
# -> ALL CHECKS PASSED

For a fast inner-loop check of the gate internals (no model, no test suite spawned), run the unit harness:

python3 plugins/fofo/skills/sdlc/scripts/selftest
# -> OK

Updating

Bump version in plugins/fofo/.claude-plugin/plugin.json and the matching entry in .claude-plugin/marketplace.json, commit, tag with claude plugin tag ./plugins/fofo, and push. Users get it via /plugin marketplace update fofo-marketplace.

Scope & status

Beta. The engine is sound and every suite passes end to end (selftest + both smokes, gated in CI and a pre-push hook); the label invites real-world use before a stable 1.0.

What it does today:

  • A test-first loop with separated test-author and code-author contexts (hard isolation on Claude Code subagents), plus tamper evidence: the Judge's tests are hash-locked at Phase 2 (test-lock) and the referee's own config is protected from the diff (oversight-integrity).
  • Deterministic Tier-0 gates on every change (spec-lint, trace-gate, redgreen-gate, intent-gate, secret-scan, test-lock, oversight-integrity, deps-gate); opt-in separation, flake, diff-budget, and model-based Tier-1 review.
  • One gate contract, so any tool — mutation, coverage, lint — drops in and routes failures the same way.
  • Spec interop: import Spec Kit / Kiro / BMAD specs via spec-adapter; spec-lint reads EARS criteria natively.

What it deliberately doesn't do yet — and where your ideas come in:

  • It's a library of gates + a procedure, not a hosted service: CI integration is one example GitHub Action, not a turnkey app.
  • intent-gate is deepest on JS/TS; other languages are token-level until a parser is contributed.
  • Separation is best-effort on hosts without context isolation.
  • No dashboard, metrics, or historical trend tracking.

Have an idea for what it should do? Open an issue or discussion →

Contributing & License

Contributions welcome — see CONTRIBUTING.md. Brand assets and usage in BRANDING.md.

MIT © 2026 Forrester Terry. Bundles acorn (MIT). See LICENSE.

About

Test integrity for AI coding agents: the context that grades the tests is isolated from the one that writes the code, and test quality is enforced by scripts with real exit codes — not prose. Claude Code plugin: /fofo:sdlc

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages