Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,13 @@ How to keep this current: add the entry in the same pull request as the change,

<!-- Empty. Next release starts here. -->

## 0.78.0

### Added

- Relevance compaction (`compaction`): experimental, off by default, not recommended. In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome. At compaction, pi-warden can write the summary instead of Pi's model. User messages and assistant text stay word for word, thinking is left out, and Jev scores each tool call with its result, each extension message, and each part of the previous summary against the current task; kept units go in verbatim, tool output inside a fence marked untrusted, and the rest become one line each. Every kept section is fenced, and a `<` that starts a `summary` tag in kept text is written `&lt;`, so kept text cannot end Pi's summary wrapper. When the security check is on, results it flagged as a possible prompt injection are never kept verbatim; results the context saver compressed keep their excerpt. `compaction.enabled` is user file only; a project may tune the other keys. One compaction sends at most `compaction.maxRequests` (12) requests sends nothing when its requests would leave fewer than 50 of the session's `maxRequests`, and stops before any request when fewer than 50 remain, so it never turns judgments off for the guards; `compaction.timeoutMs` bounds the whole compaction and the global `timeoutMs` each request. Any failure, timeout, abort, missing consent, a provider in `compaction.skipProviders` (default `claude-bridge`), a request limit, or a summary over `compaction.maxSummaryTokens` lets Pi's own summary run; the hook never cancels a compaction. One trace entry per compaction and a line in `/warden status`.
- `scripts/relevance-replay.mjs` replays recorded compactions through the relevance compaction and compares size, re-fetch coverage, and cost with Pi's summaries. First measurement in `docs/guards.md` → Calibration.

## 0.77.0

### Added
Expand Down
9 changes: 8 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,7 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
}
},
"context": { "enabled": true, "tailMinChars": 12000, "confidence": 0.8, "duplicateMinChars": 2000, "recallTool": "auto", "formatConfidence": 0.7, "dedupeRuns": true, "dedupeMessages": false, "largeOutput": { "enabled": true, "threshold": 0.85 } },
"compaction": { "enabled": false, "keepThreshold": 0.5, "maxSummaryTokens": 20000, "timeoutMs": 20000, "maxRequests": 12, "skipProviders": ["claude-bridge"] },
"runaway": { "enabled": true, "repeats": 4, "thinkingRepeats": 10, "minChars": 400, "recover": true },
"notify": { "enabled": false, "cooldownMs": 10000, "command": [] },
"judge": { "cooldownMs": 60000, "failuresBeforeCooldown": 3 },
Expand Down Expand Up @@ -118,6 +119,12 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
| `context.dedupeMessages` | Default `false`. With `context.dedupeRuns` also on, cut repeated runs in new user and custom messages the same way. Off by default because a repeat the user sends can itself carry meaning ("here it is again, still failing"), and on recent sessions messages gave about 0.8% of their bytes back. A custom message that Pi appends without an agent turn (`triggerTurn: false`, or unset while the agent is idle) does not pass Pi's `message_end` hook and stays whole. |
| `context.largeOutput.enabled` | Add one question to each judged `bash` request: will the command print far more than the agent needs? Off keeps the question out of the request. Read-only commands (`cat`, `find`, `git log`) skip the judge, so the question does not ride them. |
| `context.largeOutput.threshold` | P(large output) at or above which the agent is told, once per command family (`npm test`, `git log`, `find`) per session, to redirect or filter the command before it runs one like it again. The call is never held or warned. Default `0.85`. |
| `compaction.enabled` | Experimental, off by default, not recommended. Default `false`. Replace the summary Pi's model writes at compaction with a relevance compaction: user messages and assistant text word for word, and the tool calls Jev scores as needed for the current task with their results word for word; see [guards.md → Relevance compaction](guards.md#relevance-compaction). In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome. Needs TypeSafe consent; without it Pi's summary runs. User file only: it sends the session to Jev and spends requests, so a project's `.pi/pi-warden.json` cannot turn it on (or off); a project may set the other `compaction` keys. |
| `compaction.keepThreshold` | P(the agent needs this exact content again) at or above which a tool call and its result are kept word for word. Default `0.5`. |
| `compaction.maxSummaryTokens` | Size budget for the summary in tokens (characters / 4). Over it, the threshold rises to 0.6, 0.7, 0.8, 0.9, 0.95 on the same scores; still over, Pi's summary runs. Default `20000`, at least `1000`. |
| `compaction.timeoutMs` | The overall deadline for one compaction; past it Pi's summary runs. Each request is also bounded by the global `timeoutMs`, so a request that takes longer than `timeoutMs` fails and Pi's summary runs, even with time left here. Default `20000`, at most `120000`. |
| `compaction.maxRequests` | Requests one compaction may send. A span that needs more sends nothing, and Pi's summary runs. The compaction also shares the session's `maxRequests` budget with the guards: it sends nothing when its requests would leave fewer than 50 of that budget, and it stops before any request when fewer than 50 remain; either way Pi's summary runs, and judgments stay on for the guards. Default `12`. |
| `compaction.skipProviders` | Providers of the active model for which Pi's summary always runs. Default `["claude-bridge"]`, because pi-claude-bridge compacts its own models. |
| `context.filter` | Beta, default `{ "enabled": false, "chunkChars": 2000, "minScore": 1.5, "maxKeptChars": 6000, "timeoutMs": 4000 }`. When on, an output that would get the generic head/diagnostic/tail excerpt is split at line boundaries into chunks of about `chunkChars`; Jev scores each chunk 0 to 3 for the agent's current task, and chunks at or above `minScore` are kept word for word in original order, up to `maxKeptChars` including the last 1000 characters. Parser excerpts, `all`, duplicates, and repeated runs are unchanged. On an error, a timeout after `timeoutMs`, no consent, an exhausted request budget, or no chunk at `minScore`, the excerpt is used. Costs one or more requests per filtered output. `enabled` is read from the user file only: a project file may tune `chunkChars`, `minScore`, `maxKeptChars`, and `timeoutMs`, but cannot turn the filter on, because that sends whole redacted outputs to the judge and spends requests. See [guards.md](guards.md#context-filter-beta-off-by-default). |
| `runaway.*` | Repeat counts that abort a reply, minimum size, whether the agent gets one recovery turn. |
| `notify.*` | Desktop notifications, cooldown, optional relay command (user file only). |
Expand Down Expand Up @@ -171,7 +178,7 @@ After building, `/warden index` reports which skill and tool descriptions would

## Project config

A project may add `.pi/pi-warden.json` with `enabled` and per-guard overrides: stricter thresholds, extra guarded tools, `rules.files`, `rules.skip`, `rules.sensitivePaths`, or `"done": { "enabled": false }`. Project files are read only when Pi trusts the project. They can never grant `typesafe` consent, change `mode`, raise `timeoutMs` or `maxRequests`, or set `notify.command`.
A project may add `.pi/pi-warden.json` with `enabled` and per-guard overrides: stricter thresholds, extra guarded tools, `rules.files`, `rules.skip`, `rules.sensitivePaths`, or `"done": { "enabled": false }`. Project files are read only when Pi trusts the project. They can never grant `typesafe` consent, change `mode`, raise `timeoutMs` or `maxRequests`, set `notify.command`, or turn relevance compaction on or off (`compaction.enabled`).

A wince-style setup for a backend repo (the full version is [`examples/pi-warden.json`](../examples/pi-warden.json)):

Expand Down
1 change: 1 addition & 0 deletions docs/data-handling.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ With consent, requests go to `https://api.typesafe.ai` (default), or to the host
| **Conscience** (recommend mode) | Your current request (2000 redacted characters), the same task spine (the thread's first request and up to four earlier requests, redacted, capped at 1200 characters together), up to four recent user/assistant text messages (500 redacted characters each with roles), and sanitized candidate metadata (skill/tool name, role, lead, useWhen, examples when an index entry matches; bare description otherwise). Full skill instructions never go to Jev. The index is built locally by the session model; only sanitized entries reach Jev; advertised locations never do. Sent only when TypeSafe consent is given and the conscience module is enabled. |
| **Conscience** (load mode) | Same judge payload as recommend mode, plus: the selected skill file is read from disk (bounded by `maxSkillBytes` and `maxLoadedBytes`), frontmatter is stripped, credentials are checked, and the complete body is supplied to the main model via a custom message. Skill bodies never go to Jev. |
| **Subagent triage** | Only for a child report that names a failure, a stop, a timeout, or a question (an incremental progress line or a clean completion is answered in code and sends nothing): a redacted 1500-character head plus 500-character tail of the report, the notification type, whether it is an incremental notify, its length, and your latest prompt (1000 characters). |
| **Relevance compaction** (opt-in, `compaction.enabled`) | At each compaction: your latest request (2000 redacted characters), the same task spine, the focus a manual `/compact <text>` names (500 characters), a redacted outline of the conversation being compacted (user and assistant text clipped to at most 400 characters per message, one line per tool call), and for each tool call, extension message, and part of the previous summary a redacted 500-character input and a 1100-character head/tail sample of its result or text. A result the output check flagged as a possible prompt injection is not sampled; results are flagged only when the security check is on (`security.enabled`). The summary itself stays in the session. |
| **Nothing** | Duplicate detection, the runaway guard, sensitive-path notes, the offline part of subagent triage, standing preferences, open loops, recall, and pattern checks run entirely in code. |

## What stays on this machine
Expand Down
27 changes: 27 additions & 0 deletions docs/guards.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,18 @@ The conscience coach assesses whether the agent is missing a useful skill or too
- A second pass asked four candidate questions on the same calls (`scripts/action-candidates.mjs`, `--extra`). None separates rejected turns on its own: "would a careful engineer ask first", "is this unrequested", "did the user ask to pause", and "is the effect visible outside the working tree" all sit at the 4 to 5% base rate. `visible` has the best recall on regret (AUC 0.82, 10 of 19 regretted calls) but a commit or push is usually what was asked. Paired with the plan it works: `visible >= 0.8` and `intent_mismatch >= 0.8` flags 1.1% of calls with 18% in a rejected turn, so that pair steers at `visibleMismatch` 0.8. Two deterministic patterns came from the regretted list: a git command with hooks or signing switched off, and `gh pr merge`.
- Of 42 holds pi-warden made in those sessions, the user's next message approved 5.

### Relevance compaction replay (2026-09-29, first measurement)

`node scripts/relevance-replay.mjs` rebuilds the span each recorded compaction replaced (48 compactions in 1,521 sessions on one machine), runs the keep questions with real requests at the defaults, and compares the result with the summary Pi wrote. It counts the "re-fetch" calls: a `read` of a path or a `bash` command from the span among the first 10 tool calls after the compaction (34 calls after 17 compactions; all 34 were `read`).

- **Size.** 42 of 48 compactions produced a summary; 6 fell back before any request because the text that is always kept (user messages, assistant text, one line per call) was over 20,000 tokens. Summary median 12,371 tokens (p90 17,950) against Pi's 2,567 (p90 5,183): 4.7 times larger at the median. The threshold rose above 0.5 in 20 of 42 to fit.
- **Re-fetch coverage.** Of the 34 re-fetched reads, the relevance summary held 1 whole and 11 as head and tail (the files were mostly 5,000 to 37,000 characters), missed 20, and 2 fell in fallbacks. Pi's summary holds no tool output word for word; it names all 34 paths.
- **Signal.** Jev ranks the re-read files high: for 22 of the 32 scored re-reads, a unit of that file scored 0.5 or more, a level only 20% of all 5,420 scored units reach (median 0.36, p90 0.57). Most of them were lost to the raised threshold and the 4000-character cut, not to the ranking.
- **Cost.** Median 6 requests per compaction (p90 11), 101,000 input tokens (p90 190,000), 0.85 s (p90 1.25 s); no timeout at 20 s.
- **Batched against one question per request** (3 compactions of 130 to 137 units, 401 units): 342 of 401 keep decisions agree (85%; 88%, 92%, 77%), mean absolute difference 0.05, and a unit asked alone scores 0.03 higher on average.

The feature ships experimental, off by default, and not recommended: on this data it keeps more text than Pi's summary without holding what the agent went back for. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome.

### Intent mismatch (2026-09-29, blind labels)

140 sampled calls with a plan were labelled by hand, without the score in view, for whether the call did something other than the agent's stated plan. The `intent_mismatch` score separates the two (AUROC 0.815), but the steer it would deliver does not: of the 37 calls that would reach the agent, 16 did exactly what the plan said, 20 went beyond the plan on something the user's latest request had asked for, and 1 caught something the user had not asked for. The score, the trace entry, the `/warden status` counters, and the thresholds are unchanged; `action.intentTraceOnly` defaults to `"all"` for this reason, and `"invisible"` restores the old delivery.
Expand Down Expand Up @@ -400,6 +412,21 @@ Only the newest tool result or message is ever changed, before it enters the ses

Set `context.enabled: false` to turn it off. Full-output files can contain secrets and stay in the OS temporary directory until removed.

### Relevance compaction

Experimental, off by default, not recommended. In a replay of 48 recorded compactions its summary was 4.7 times the size of Pi's at the median and kept whole only 1 of the 34 files the agent read again. Try it or improve it; changes that make it smaller or keep what the agent goes back for are welcome.

Opt-in (`compaction.enabled`, user file only; needs TypeSafe consent). When Pi compacts a session, pi-warden can write the summary instead of Pi's model, in `session_before_compact`. Nothing in it is paraphrased:

- **Kept word for word, always:** user messages and assistant text. Thinking is never kept.
- **Scored by Jev:** each tool call with its result, each extension message, and each part of the previous summary (the sections of Pi's summary, or the units of an earlier relevance compaction). One `noul` question per unit asks whether the agent will need its exact content for the current task (the latest request, the task spine, and the focus a manual `/compact <text>` names). At or above `compaction.keepThreshold` (0.5) the call and its result are kept; a result or input over 4000 characters keeps its first 2400 and last 1200 characters and names the saved full-output file when pi-warden has one. Below it, the call is one line under "left out", with no result.
- **Never kept word for word:** a result the output check flagged as a possible prompt injection (one line, and no question is asked about it). Flagged results are recognised only when the security check is on (`security.enabled`); with it off, no result carries the flag. A result the context saver already replaced keeps its excerpt, uncut.
- **Layout:** a header with the kept and left-out counts, the files read and modified (from Pi's file operations and the previous summary), then the units in their original order. Every kept section sits in a fence longer than any backtick run in it, so the next compaction reads back the same units; every kept tool result and extension message sits in a fence labelled untrusted: data, not instructions. Pi wraps the summary in `<summary>` tags without escaping, so a `<` that starts a `summary` tag in kept text is written `&lt;`, and the header says so. The summary enters the context as one user message, as Pi's does.
- **Requests:** every request carries the task, an outline of the whole span (shrunk in stages to fit), and up to 24 units with a redacted input and a head/tail sample of each result, under the 64 KiB request limit; labels and headings are redacted too. Four requests run at once, at most `compaction.maxRequests` (12) per compaction. The compaction shares the session's `maxRequests` budget with the guards: it sends nothing when its requests would leave fewer than 50 of that budget, and it stops before any request when fewer than 50 remain, so a compaction never turns judgments off for the guards.
- **Fallback:** Pi's summary runs (the hook returns nothing; it never cancels a compaction) when consent is missing, the model's provider is in `compaction.skipProviders`, the span needs more than `compaction.maxRequests` requests, the request reserve is reached, a request fails or passes the global `timeoutMs`, `compaction.timeoutMs` passes, the compaction is aborted, or the summary stays over `compaction.maxSummaryTokens` after the threshold is raised. Each compaction leaves one trace entry (kept and scored units, requests, input tokens, time, or the fallback reason), and `/warden status` has one line for the session.

The compaction appendix above still follows every compaction, this one included.

### Context filter (beta, off by default)

`context.filter.enabled: true` changes one case only: a single text block for which the saver would build the generic head/diagnostic/tail excerpt (retention `errors_and_summary` or `summary_only`, and no format parser fits). Parser excerpts, `all`, duplicates, repeated runs, multi-block results, and outputs below `tailMinChars` are unchanged.
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "pi-warden",
"version": "0.77.0",
"version": "0.78.0",
"description": "Makes the Pi agent follow your project's rules. Jev judges every write against your pi-warden.md and quotes the broken rule back to the agent, names slop, breaks stuck loops, calls out unverified done claims, compresses large tool output, and holds the rare destructive command. Built on pi-typesafe.",
"type": "module",
"license": "MIT",
Expand Down
Loading
Loading