Your coding agent says "done". It didn't run the tests.
pi-warden catches it, and the agent fixes it without you.
759 sessions · 220 risky actions stopped before they ran · 76% of fake "done"s turned into real test runs
Nine days of the maintainer's real use across all their projects, not just this one, 2026-09-16 to 2026-09-24. Method and noise: field report.
pi install npm:pi-wardenThen /warden enable (paste a TypeSafe key) and /warden init (writes a starter pi-warden.md). That's it. No key? The offline guards still run.
Most guardrails stop and ask you. pi-warden tells the agent what it got wrong, and the agent corrects itself. You are pulled in only when something can't be undone: 3 holds per 1,000 calls. The other 997 just run.
| Your agent… | pi-warden… |
|---|---|
| says "done" with no test, build, or lint behind it | sends it back to prove it |
breaks a rule in your pi-warden.md or AGENTS.md |
quotes the exact rule it broke |
is about to git push --force, reset --hard, rm -rf, DROP |
holds it before it runs |
| retries the same failing fix for the third time | asks for a new hypothesis |
| writes stubs, restating comments, hardcoded secrets | names them on the spot |
| floods its context with a 40k-line log | keeps the lines that matter, stores the rest |
| starts repeating itself forever | stops the reply |
Every guard, with its thresholds and calibration →
# A TODO names a ticket
A bare `TODO` or `FIXME` without a ticket reference is a violation.
# Errors never reach the user raw
paths: src/api/**/*.ts
Catch errors at the handler and return a message a user can act on.Each # heading is one rule. Every write and edit is judged against it in about a quarter of a second by Jev, a model that returns a probability, not prose. No pi-warden.md? Your AGENTS.md, CLAUDE.md, or README.md is used instead.
- Rules: in 150 paired agent runs, the agent without pi-warden broke the tested rule 6 times. With it: 0.
- Done-check: after a nudge, the agent ran a check 57 of 75 times, and sometimes found a failure it had missed.
- Holds: when the agent was stopped, it found a safer way 40 of 65 times; you approved 24.
- Stability: 13,952 guard cases over 109 overnight cycles, no score drift.
Every number has a script and a raw report in eval/reports/. They are the maintainer's measurements, not a universal promise, and the reports list what was noise.
Secrets and unshown paths are stripped before anything leaves your machine. Exactly what is sent →
Guards · Configuration · Commands · FAQ · Data handling · Examples
npm install
npm run check # typecheck + offline tests + buildMIT licensed.
