Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,13 @@ How to keep this current: add the entry in the same pull request as the change,

<!-- Empty. Next release starts here. -->

## 0.78.0

### Added

- Relevance compaction (`compaction`): experimental, off by default, not recommended. In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome. At compaction, pi-warden can write the summary instead of Pi's model. User messages and assistant text stay word for word, thinking is left out, and Jev scores each tool call with its result, each extension message, and each part of the previous summary against the current task; kept units go in verbatim, tool output inside a fence marked untrusted, and the rest become one line each. Every kept section is fenced, and a `<` that starts a `summary` tag in kept text is written `&lt;`, so kept text cannot end Pi's summary wrapper. When the security check is on, results it flagged as a possible prompt injection are never kept verbatim; results the context saver compressed keep their excerpt. `compaction.enabled` is user file only; a project may tune the other keys. One compaction sends at most `compaction.maxRequests` (12) requests sends nothing when its requests would leave fewer than 50 of the session's `maxRequests`, and stops before any request when fewer than 50 remain, so it never turns judgments off for the guards; `compaction.timeoutMs` bounds the whole compaction and the global `timeoutMs` each request. Any failure, timeout, abort, missing consent, a provider in `compaction.skipProviders` (default `claude-bridge`), a request limit, or a summary over `compaction.maxSummaryTokens` lets Pi's own summary run; the hook never cancels a compaction. One trace entry per compaction and a line in `/warden status`.
- `scripts/relevance-replay.mjs` replays recorded compactions through the relevance compaction and compares size, re-fetch coverage, and cost with Pi's summaries. First measurement in `docs/guards.md` → Calibration.

## 0.77.0

### Added
Expand Down
67 changes: 61 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,20 +34,61 @@ Then `/warden enable` (paste a [TypeSafe](https://console.typesafe.ai) key) and

## It steers. It doesn't nag.

Most guardrails stop and ask you. pi-warden tells **the agent** what it got wrong, and the agent corrects itself. You are pulled in only when something can't be undone: **3 holds per 1,000 calls. The other 997 just run.**
Most guardrails stop and ask you. pi-warden tells **the agent** what it got wrong, and the agent corrects itself. You are pulled in only when something can't be undone. In a replay of 18,075 recorded calls (2026-09-21), **0.27% were held at the old 0.7 threshold and 0.1% at 0.9**, the default since 0.75.0 ([report](eval/reports/2026-09-21-calibration-0.33.3/report.md)).

| Your agent… | pi-warden… |
| --- | --- |
| says "done" with no test, build, or lint behind it | sends it back to prove it |
| breaks a rule in your `pi-warden.md` or `AGENTS.md` | quotes the exact rule it broke |
| is about to `git push --force`, `reset --hard`, `rm -rf`, `DROP` | holds it before it runs |
| is about to `git push --force`, `reset --hard`, `rm -rf`, `DROP` | holds it before it runs when the judge is at least 0.9 sure it can't be undone; with no judge, a built-in pattern holds it |
| retries the same failing fix for the third time | asks for a new hypothesis |
| writes stubs, restating comments, hardcoded secrets | names them on the spot |
| floods its context with a 40k-line log | keeps the lines that matter, stores the rest |
| floods its context with a huge log | keeps the lines that matter, stores the rest |
| starts repeating itself forever | stops the reply |

[Every guard, with its thresholds and calibration →](docs/guards.md)

## Features

Every guard and feature, its default, and its status. Defaults are `defaultConfig()` in `src/config.ts`; Beta and Experimental are the labels in [the guard docs](docs/guards.md). Guards that ask Jev need a key and `/warden enable`; without one, their offline parts still run.

| Feature | What it does | Default | Status |
| --- | --- | --- | --- |
| Jev judgments | Sends redacted samples to Jev for the judged checks below | Off until `/warden enable` | Stable |
| [Action guard](docs/guards.md#action-guard) | Checks each command, write, and edit before it runs: offline patterns plus one Jev request | On | Stable |
| ↳ Irreversible hold | Holds a call Jev scores 0.9 or more as irreversible; 0.5 to 0.9 warns | On | Stable |
| ↳ Off-task steer | Warns at 0.6 and steers the agent back to your request at 0.85; never holds | On | Stable |
| ↳ Intent-mismatch steer | Tells the agent when a call differs from its stated plan; the score stays in the trace | Off: trace only. `action.intentTraceOnly: "invisible"` | Stable |
| ↳ Should-proceed steer | Asks the agent to pause and ask you; the score stays in the trace | Off: trace only. `action.shouldProceed.steer: true` | Stable |
| ↳ Your command, path, and arming rules | Warn, hold, or deny commands and paths you name | On, none set | Stable |
| ↳ Large-output warning | Tells the agent to filter a command that will print far more than it needs | On | Stable |
| [Rules](docs/guards.md#rules) | Judges every write and edit against the rules in your `pi-warden.md` | On | Stable |
| ↳ Soft double-check tier | Asks the agent to double-check a score just below a rule's cutoff | Off. `rules.softThreshold` | Stable |
| [Slop](docs/guards.md#slop) | Names stubs, restating comments, dead code, and padded replies | On | Stable |
| [Security](docs/guards.md#security) | Flags risky written code and prompt injection in tool output; masks credentials | On | Stable |
| [Stuck](docs/guards.md#stuck) | Asks for a new hypothesis when the agent repeats a failing approach | On | Stable |
| [Done-check](docs/guards.md#done-check) | Sends an unverified "done" back for a check; after a UI change, a visual check | On | Stable |
| [Runaway](docs/guards.md#runaway) | Stops a reply that repeats itself, offline | On | Stable |
| [Context saver](docs/guards.md#context-saver) | Replaces large or duplicate tool output with an excerpt and a saved full copy | On | Stable |
| ↳ Repeated user messages | Cuts repeated runs in your own messages too | Off. `context.dedupeMessages` | Stable |
| ↳ Compaction evidence appendix | After a compaction, lists failed calls, the last passing check, holds, and saved outputs | On | Stable |
| [Context filter](docs/guards.md#context-filter-beta-off-by-default) | Jev scores chunks of a large output and keeps the ones for the current task | Off. `context.filter.enabled` | Beta |
| [Relevance compaction](docs/guards.md#relevance-compaction) | Writes the compaction summary instead of Pi's model; not recommended | Off. `compaction.enabled` (user file only) | Experimental |
| [Call-waste notes](docs/guards.md#call-waste) | One advisory line on polling, paging, repeated searches, and re-filtered checks | On | Stable |
| ↳ Session tip | Adds one paragraph about call cost to the system prompt | Off. `waste.tip` | Stable |
| [Open loops and recall](docs/guards.md#open-loops-and-recall) | `warden_loops` keeps the agent's promises; `warden_recall` lists what it already tried | On, no switch | Stable |
| [Subagent triage](docs/guards.md#subagent-triage) | Wakes the agent only for a subagent report that needs it | On | Stable |
| [Judge cooldown](docs/guards.md#judge-cooldown) | Pauses Jev requests after repeated failures, so a dead backend costs no timeouts | On, no switch | Stable |
| [Steer budget](docs/guards.md#steer-messages) | At most 3 non-critical steers per run; the rest go to the trace | On | Stable |
| [Adaptive steers](docs/guards.md#adaptive-steers-per-model) | Makes a steer kind trace-only for a model that rarely follows or often disputes it | On | Stable |
| Steers and notices in the transcript | Shows steers and per-call notices in the chat, not only in the trace | Off. `steerVisible`, `notices` | Stable |
| [Desktop notifications](docs/guards.md#desktop-notifications) | Holds, confirm dialogs, and runaway stops reach the desktop or your relay | Off. `notify.enabled` | Stable |
| [Conscience](docs/guards.md#conscience) | Recommends a skill or tool the agent is missing | Off. `conscience.enabled` | Beta |
| Learning from holds | Records hold outcomes locally; `/warden recommend` suggests threshold changes | On | Stable |
| Standing preferences | `/warden prefs` finds preferences you repeated across sessions and sends them at session start | On | Stable |

All keys and their values: [configuration](docs/configuration.md). Commands: [commands](docs/commands.md).

## Rules no linter can check

```markdown
Expand All @@ -65,14 +106,28 @@ Each `#` heading is one rule. Every write and edit is judged against it in about

- **Rules:** in 150 paired agent runs, the agent without pi-warden broke the tested rule **6 times**. With it: **0**.
- **Done-check:** after a nudge, the agent ran a check **57 of 75** times, and sometimes found a failure it had missed.
- **Holds:** when the agent was stopped, it found a safer way **40 of 65** times; you approved 24.
- **Holds:** when the agent was stopped, it found a safer way **40 of 65** times; you approved 24. Field data from 2026-09-16 to 2026-09-24, before the 0.9 threshold.
- **Stability:** 13,952 guard cases over 109 overnight cycles, no score drift.

Every number has a script and a raw report in [`eval/reports/`](eval/reports/). They are the maintainer's measurements, not a universal promise, and the reports list what was noise.
Every number has a script and a raw report in [`eval/reports/`](eval/reports/) or the Calibration section of [the guard docs](docs/guards.md#calibration). They are the maintainer's measurements, not a universal promise, and the reports list what was noise.

## How the defaults are chosen

Thresholds are set on recorded sessions. Where the data allows, a threshold is chosen on one half of the corpus and checked on the other. A feature that measured worse ships off or trace-only. [Calibration →](docs/guards.md#calibration)

- **Irreversible hold at 0.9** (0.75.0). On 15,346 judged calls the judge was wrong on 15% of calls below confidence 0.8 and under 1% above it. A 0.9 cutoff chosen on one half removed about 52 false alarms on the other half and lost no true catch. Scores from 0.5 to 0.9 now warn.
- **Intent steer in the trace only** (0.76.0). The score separates a differing call (AUROC 0.815 on 140 hand-labelled calls), but 36 of the 37 steers it would send were for calls the plan or your request had asked for.
- **Relevance compaction off** (0.78.0). In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again.

## Privacy

Secrets and unshown paths are stripped before anything leaves your machine. [Exactly what is sent →](docs/data-handling.md)
Secrets and unshown paths are stripped before anything leaves your machine. The offline guards send nothing. The opt-in features send more when you turn them on: the context filter sends a large output whole, redacted, in chunks; relevance compaction sends an outline of the conversation and a sample of each tool result. [Exactly what is sent →](docs/data-handling.md)

## Credits

- **Confidence bands for the irreversible hold:** Li, Miao, Krishnan, Padman, "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure", [arXiv:2609.26550](https://arxiv.org/abs/2609.26550).
- **Chunk scoring in the context filter:** GPT Researcher's [context filter](https://docs.gptr.dev/docs/gpt-researcher/gptr/context-filter).
- **The judge:** [Jev](https://typesafe.ai) by TypeSafe.

## Docs

Expand Down
9 changes: 8 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,7 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
}
},
"context": { "enabled": true, "tailMinChars": 12000, "confidence": 0.8, "duplicateMinChars": 2000, "recallTool": "auto", "formatConfidence": 0.7, "dedupeRuns": true, "dedupeMessages": false, "largeOutput": { "enabled": true, "threshold": 0.85 } },
"compaction": { "enabled": false, "keepThreshold": 0.5, "maxSummaryTokens": 20000, "timeoutMs": 20000, "maxRequests": 12, "skipProviders": ["claude-bridge"] },
"runaway": { "enabled": true, "repeats": 4, "thinkingRepeats": 10, "minChars": 400, "recover": true },
"notify": { "enabled": false, "cooldownMs": 10000, "command": [] },
"judge": { "cooldownMs": 60000, "failuresBeforeCooldown": 3 },
Expand Down Expand Up @@ -118,6 +119,12 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
| `context.dedupeMessages` | Default `false`. With `context.dedupeRuns` also on, cut repeated runs in new user and custom messages the same way. Off by default because a repeat the user sends can itself carry meaning ("here it is again, still failing"), and on recent sessions messages gave about 0.8% of their bytes back. A custom message that Pi appends without an agent turn (`triggerTurn: false`, or unset while the agent is idle) does not pass Pi's `message_end` hook and stays whole. |
| `context.largeOutput.enabled` | Add one question to each judged `bash` request: will the command print far more than the agent needs? Off keeps the question out of the request. Read-only commands (`cat`, `find`, `git log`) skip the judge, so the question does not ride them. |
| `context.largeOutput.threshold` | P(large output) at or above which the agent is told, once per command family (`npm test`, `git log`, `find`) per session, to redirect or filter the command before it runs one like it again. The call is never held or warned. Default `0.85`. |
| `compaction.enabled` | Experimental, off by default, not recommended. Default `false`. Replace the summary Pi's model writes at compaction with a relevance compaction: user messages and assistant text word for word, and the tool calls Jev scores as needed for the current task with their results word for word; see [guards.md → Relevance compaction](guards.md#relevance-compaction). In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome. Needs TypeSafe consent; without it Pi's summary runs. User file only: it sends the session to Jev and spends requests, so a project's `.pi/pi-warden.json` cannot turn it on (or off); a project may set the other `compaction` keys. |
| `compaction.keepThreshold` | P(the agent needs this exact content again) at or above which a tool call and its result are kept word for word. Default `0.5`. |
| `compaction.maxSummaryTokens` | Size budget for the summary in tokens (characters / 4). Over it, the threshold rises to 0.6, 0.7, 0.8, 0.9, 0.95 on the same scores; still over, Pi's summary runs. Default `20000`, at least `1000`. |
| `compaction.timeoutMs` | The overall deadline for one compaction; past it Pi's summary runs. Each request is also bounded by the global `timeoutMs`, so a request that takes longer than `timeoutMs` fails and Pi's summary runs, even with time left here. Default `20000`, at most `120000`. |
| `compaction.maxRequests` | Requests one compaction may send. A span that needs more sends nothing, and Pi's summary runs. The compaction also shares the session's `maxRequests` budget with the guards: it sends nothing when its requests would leave fewer than 50 of that budget, and it stops before any request when fewer than 50 remain; either way Pi's summary runs, and judgments stay on for the guards. Default `12`. |
| `compaction.skipProviders` | Providers of the active model for which Pi's summary always runs. Default `["claude-bridge"]`, because pi-claude-bridge compacts its own models. |
| `context.filter` | Beta, default `{ "enabled": false, "chunkChars": 2000, "minScore": 1.5, "maxKeptChars": 6000, "timeoutMs": 4000 }`. When on, an output that would get the generic head/diagnostic/tail excerpt is split at line boundaries into chunks of about `chunkChars`; Jev scores each chunk 0 to 3 for the agent's current task, and chunks at or above `minScore` are kept word for word in original order, up to `maxKeptChars` including the last 1000 characters. Parser excerpts, `all`, duplicates, and repeated runs are unchanged. On an error, a timeout after `timeoutMs`, no consent, an exhausted request budget, or no chunk at `minScore`, the excerpt is used. Costs one or more requests per filtered output. `enabled` is read from the user file only: a project file may tune `chunkChars`, `minScore`, `maxKeptChars`, and `timeoutMs`, but cannot turn the filter on, because that sends whole redacted outputs to the judge and spends requests. See [guards.md](guards.md#context-filter-beta-off-by-default). |
| `runaway.*` | Repeat counts that abort a reply, minimum size, whether the agent gets one recovery turn. |
| `notify.*` | Desktop notifications, cooldown, optional relay command (user file only). |
Expand Down Expand Up @@ -171,7 +178,7 @@ After building, `/warden index` reports which skill and tool descriptions would

## Project config

A project may add `.pi/pi-warden.json` with `enabled` and per-guard overrides: stricter thresholds, extra guarded tools, `rules.files`, `rules.skip`, `rules.sensitivePaths`, or `"done": { "enabled": false }`. Project files are read only when Pi trusts the project. They can never grant `typesafe` consent, change `mode`, raise `timeoutMs` or `maxRequests`, or set `notify.command`.
A project may add `.pi/pi-warden.json` with `enabled` and per-guard overrides: stricter thresholds, extra guarded tools, `rules.files`, `rules.skip`, `rules.sensitivePaths`, or `"done": { "enabled": false }`. Project files are read only when Pi trusts the project. They can never grant `typesafe` consent, change `mode`, raise `timeoutMs` or `maxRequests`, set `notify.command`, or turn relevance compaction on or off (`compaction.enabled`).

A wince-style setup for a backend repo (the full version is [`examples/pi-warden.json`](../examples/pi-warden.json)):

Expand Down
Loading
Loading