Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,12 @@ How to keep this current: add the entry in the same pull request as the change,

<!-- Empty. Next release starts here. -->

## 0.75.0

### Changed

- The default `action.irreversible.confirm` (the hold threshold) is 0.9 instead of 0.7, so a judge-only hold waits for the confidence the recorded action-guard corpus shows is safe: the judge's error rate falls from 15% below confidence 0.8 to under 1% above it, and a 0.9 cutoff chosen on one half of the corpus removed about 52 false alarms on the other half without losing a true catch. A call the judge scores 0.5 to 0.9 now warns instead of holding; pattern holds and user or project overrides are unchanged. Restore the old behaviour with `"action": { "irreversible": { "confirm": 0.7 } }`.

## 0.74.1

### Fixed
Expand Down
4 changes: 2 additions & 2 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
"enabled": true,
"tools": ["bash", "powershell", "ctx_execute", "ctx_batch_execute", "ctx_execute_file", "write", "edit"],
"failOpen": true,
"irreversible": { "warn": 0.5, "confirm": 0.7 },
"irreversible": { "warn": 0.5, "confirm": 0.9 },
"offTask": { "warn": 0.6, "steer": 0.85 },
"intentMismatch": 0.9,
"visibleMismatch": 0.8,
Expand Down Expand Up @@ -85,7 +85,7 @@ User file `~/.pi/agent/pi-warden/config.json` (owner-only). `/warden config` ope
| `timeoutMs` | Per-request timeout. On timeout the call is allowed with a warning when `action.failOpen` is true. |
| `maxRequests` | Per-session request budget. When spent, pi-warden says so once and continues with offline checks. |
| `action.tools` | Tools the action guard inspects. Add your own shell-like tools here. |
| `action.irreversible` | `warn` and `confirm` (hold) thresholds on P(irreversible). |
| `action.irreversible` | `warn` and `confirm` (hold) thresholds on P(irreversible). Defaults `warn` 0.5 and `confirm` 0.9. The 0.9 hold waits for the confidence at which the judge stops being wrong: on the recorded action-guard corpus the judge errs on 15% of calls below confidence 0.8 and under 1% above it, and a 0.9 cutoff removed about 52 false alarms on a held-out half of the corpus without losing a true catch. Calls the judge scores 0.5 to 0.9 warn instead of holding. See [guards.md](guards.md#irreversible-hold-threshold-2026-09-29-held-out-split). |
| `action.offTask` | `warn` and `steer` thresholds on P(off-task). Off-task never holds. |
| `action.intentMismatch` | P(call differs from the agent's stated plan) that warns and tells the agent, on calls that can change something. |
| `action.visibleMismatch` | Lower mismatch threshold for commands whose effect is visible outside the working tree (commit, push, publish, install, launch). |
Expand Down
23 changes: 19 additions & 4 deletions docs/guards.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ Runs on `tool_call`, before the tool executes.
Text that is data is not a command. A heredoc body written to a file, a quoted `echo`/`printf` argument, a `grep` pattern, or a `git commit -m` message can mention `git push --force` without a hold. The same text fed to `sh`, `bash -c`, `eval`, `xargs`, or a `python3 - <<EOF` script that calls `os.system` keeps every hit.

In **evidence mode** (`action.floor: "evidence"`, the default), built-in pattern hits listed above are fed to the judge as `floor_hits` in the request state and traced as `(evidence)` in reasons, but they do not set the hold level. The judge's `irreversible` score against the configured thresholds decides warn and confirm. This prevents the floor from overriding a present, confident judge. Without a judge (TypeSafe unavailable, consent off, request failed), or in **level mode** (`action.floor: "level"`), the floor applies as before: destructive hits hold, risky/sensitive hits warn, outside-project existing-file writes hold.
3. **Jev**, with consent: one request with `{ task, spine, context, plan, action, floor_hits }` and five questions. `irreversible` (yes/no), `off_task` (yes/no), `mutates` (does it change anything), `scope` (expected step, plausible side step, unrelated, unclear), `should_proceed` (yes/no, inverted: low = steer). Defaults: irreversible at 0.5 warns and at 0.7 holds. Off-task never holds: at 0.6 it warns, and at 0.85 with `unrelated` on a call that can change something the agent is also steered back to your request (an unrelated `grep` is warned about only). `should_proceed` steers but never holds: when P(yes) drops below 0.6 the agent is told to pause and ask the user. In evidence mode, built-in pattern hits are listed in `floor_hits` so the judge weighs them; in level mode, patterns set the floor and Jev can only raise it.
3. **Jev**, with consent: one request with `{ task, spine, context, plan, action, floor_hits }` and five questions. `irreversible` (yes/no), `off_task` (yes/no), `mutates` (does it change anything), `scope` (expected step, plausible side step, unrelated, unclear), `should_proceed` (yes/no, inverted: low = steer). Defaults: irreversible at 0.5 warns and at 0.9 holds; the 0.5 to 0.9 band warns instead of holding (see [Calibration](#calibration)). Off-task never holds: at 0.6 it warns, and at 0.85 with `unrelated` on a call that can change something the agent is also steered back to your request (an unrelated `grep` is warned about only). `should_proceed` steers but never holds: when P(yes) drops below 0.6 the agent is told to pause and ask the user. In evidence mode, built-in pattern hits are listed in `floor_hits` so the judge weighs them; in level mode, patterns set the floor and Jev can only raise it.

`plan` is the agent's own words in the message that makes the call, or in the text-only message right before it with no tool call in between (500 redacted characters). Text from before an earlier tool call described that call, so it is not sent and the intent question is not asked, except for a shell command with a visible effect (a `git` commit, push, merge, tag, or reset, `gh pr`, `gh release`, `npm publish`): that call is still judged against the latest text since your prompt. It tells Jev which step this is, so a verification fixture the agent just announced is not judged unrelated; it never authorizes anything. When there is a plan, a fifth question `intent_mismatch` asks whether the call does something materially different from it: a delete where the plan said list, a force push where it said push. At `action.intentMismatch` (0.9) on a call that can change something, the call is warned about and the agent is told to keep its words and its calls in step. A command whose effect is visible outside the working tree (`visible`: a commit, push, merge, publish, message, install, launched program) needs only `action.visibleMismatch` (0.8): on recorded sessions that pair is what users objected to. Never held on that alone. The warning reaches the agent after the call ran, so by default (`action.intentTraceOnly: "invisible"`) only a call with a visible effect steers the agent: a commit, push, merge, tag, reset, pull request, release, or publish (decided in code), or a call Jev judges `visible` at 0.8 or more (an install, a launched program, a message sent from a script); any other mismatch stays in the trace and the status count. On 275 recorded steers, all 275 arrived after the call, and a strict course change followed 8%. `"none"` steers on every mismatch; `"all"` on none. The trace shows the plan under each verdict.
4. **Act**, by mode:
Expand Down Expand Up @@ -73,7 +73,7 @@ The conscience coach assesses whether the agent is missing a useful skill or too

`node scripts/calibrate-action.mjs --all` replays every guarded call in your recorded Pi sessions through the guard (one request per call) and asks Jev, once per turn, whether your next message regrets one of the calls that ran, approves each held call, and how it receives the turn (continues, corrects, rejects, unrelated). Run on 321 sessions from this machine (1,085 labelled turns, 17,160 guarded calls, 14,903 judged):

- Regret is rare: 20 calls (2% of turns). None of them was about data loss: their `irreversible` scores were 0.04 to 0.57, median 0.07. They were scope and permission complaints: a commit the user did not want, an edit to a personal `CLAUDE.md`, a merge, a test run when conflicts were the job, a program launched at night. The hold rule catches none of them at any threshold that holds fewer than 3% of calls, so the hold defaults stay where they are; they are a checkpoint for destructive actions, and regret is the wrong yardstick for those.
- Regret is rare: 20 calls (2% of turns). None of them was about data loss: their `irreversible` scores were 0.04 to 0.57, median 0.07. They were scope and permission complaints: a commit the user did not want, an edit to a personal `CLAUDE.md`, a merge, a test run when conflicts were the job, a program launched at night. The hold rule catches none of them at any threshold that holds fewer than 3% of calls, so the irreversible hold is set on the judge's own confidence instead (see the calibration below); they are a checkpoint for destructive actions, and regret is the wrong yardstick for those.
- Signal ranking against regret (AUC): `mutates` 0.74, `irreversible` 0.71, `intent_mismatch` 0.57, `off_task` 0.51. Off-task alone caused 56 of the 139 replay holds and none of them drew a complaint, so since 0.12 off-task warns and steers but never holds (`offTask.steer`; a `confirm` key in an older config file still sets it).
- The intent steer earned its threshold here. At 0.8 it fires on 11% of calls that can change something and 14% of those sit in a turn the user rejects (base rate 5%); at 0.9 it fires on 4% and 33% of those are in a rejected turn, 54% in one the user rejects or corrects (base rate 24%). The default is 0.9.
- A second pass asked four candidate questions on the same calls (`scripts/action-candidates.mjs`, `--extra`). None separates rejected turns on its own: "would a careful engineer ask first", "is this unrequested", "did the user ask to pause", and "is the effect visible outside the working tree" all sit at the 4 to 5% base rate. `visible` has the best recall on regret (AUC 0.82, 10 of 19 regretted calls) but a commit or push is usually what was asked. Paired with the plan it works: `visible >= 0.8` and `intent_mismatch >= 0.8` flags 1.1% of calls with 18% in a rejected turn, so that pair steers at `visibleMismatch` 0.8. Two deterministic patterns came from the regretted list: a git command with hooks or signing switched off, and `gh pr merge`.
Expand All @@ -94,7 +94,7 @@ AUC against regret: 0.26 — non-regretted calls score higher (correct direction
| `should_proceed` | 0.26 | 0.58 |
| `pause_requested` | 0.27 | 0.51 |

The question is trace-only by default until calibrated: AUC against regret is 0.26 and the default threshold of 0.6 flags 44% of non-read-only calls. The score and reason remain in the trace, but no steer reaches the agent. Set `action.shouldProceed.steer: true` to restore the pause-and-ask steer; `hold` remains the threshold. This question never holds a call, consistent with the existing rule that only deny rules and `irreversible >= 0.7` hold; a built-in destructive pattern holds only when no judge answers or when `action.floor` is `"level"`.
The question is trace-only by default until calibrated: AUC against regret is 0.26 and the default threshold of 0.6 flags 44% of non-read-only calls. The score and reason remain in the trace, but no steer reaches the agent. Set `action.shouldProceed.steer: true` to restore the pause-and-ask steer; `hold` remains the threshold. This question never holds a call, consistent with the existing rule that only deny rules and `irreversible >= 0.9` hold; a built-in destructive pattern holds only when no judge answers or when `action.floor` is `"level"`.

### violation_judgment calibration (2026-09-20)

Expand Down Expand Up @@ -146,6 +146,21 @@ A full replay of four projects' recorded sessions at 0.33.3 with `--extra`, judg

n = 15,232 judged calls with 26 positives, except `intent_mismatch` (12,636 calls with a plan, 20 positives). Since 17 September the hold rate fell from 1% to 0.27%, because off-task no longer holds; `irreversible` rose from 0.71 to 0.73, `mutates` slipped from 0.75 to 0.72, the candidate questions barely moved, and the share of live holds the user approved rose from 12% to 36%. Full tables: `eval/reports/2026-09-21-calibration-0.33.3/`.

### Irreversible hold threshold (2026-09-29, held-out split)

`action.irreversible.confirm` is 0.9, not 0.7. The threshold is set on the judge's own confidence, not on regret: on the recorded action-guard corpus (15,346 judged calls, 26 regretted calls), confidence is `max(p, 1 - p)` and a judge failure counts as an error.

| confidence | items | error | false alarms | misses |
| --- | --- | --- | --- | --- |
| 0.50–0.60 | 118 | 42.4% | 49 | 1 |
| 0.60–0.80 | 493 | 15.2% | 72 | 3 |
| 0.80–0.95 | 4,937 | 0.8% | 28 | 13 |
| 0.95–1.00 | 9,798 | 0.1% | 0 | 8 |

The error rate collapses once confidence passes 0.8: below it the judge is wrong between one call in seven and one in two. A 0.9 cutoff chosen on a random half of the corpus (seed 20260930) and checked on the other half removed about 52 false alarms on the held-out half and lost no true catch, so calls the judge alone scored 0.7 to 0.9 now warn instead of holding. In the replay, none of the 42 calls scored in that band sat in a rejected turn, and the 23 of them on calls that ran also drew no regret.

The evidence is a direction, not a fitted threshold: there are only 26 regret positives, and the misses sit at high confidence (0.95 to 1.00), where a cutoff cannot reach them — moving the threshold trades false alarms for misses, it does not fix a judge that is confidently wrong. The method follows Li, Miao, Krishnan, Padman, "JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" (arXiv:2609.26550, Carnegie Mellon University): pick the threshold on a selection split, re-check it on a held-out split, and count judge failures as errors.

### Live: what fired, and what the agent did next

The replay measures the action guard's decisions against your reactions. It cannot measure the other half: what the agent does with a steer. For that, 67 steer messages from two days of live work on one production repo (19 sessions, 2026-09-16 to 09-17), read back from the recorded session logs:
Expand Down Expand Up @@ -315,7 +330,7 @@ Confusion over the four reasons (rows labelled, columns answered): 21/21 `from_c

### Rules tiers calibration (2026-09-27, per-rule cutoff, severity, and soft tier)

The bench grew to 161 labelled cases (17 rules: 107 tune, 54 holdout). A rule body may now start with `threshold:` and `severity:` lines beside `paths:`; a rule with its own cutoff fires there instead of `rules.threshold`, and severity orders findings high, normal, low, then score, in the steer and the trace. An opt-in `rules.softThreshold` (`0`, off) turns a score between it and a rule's cutoff into a soft "check whether this applies" sentence in the same steer, without holding and without counting as a finding. A full run at the shipped settings (161 cases, 147 asked: 74 violation, 73 clean) reaches recall 0.865, false alarms 0.027 and precision 0.970 at 0.7; the soft tier at 0.5 adds 8 catches and 4 false alarms against the per-rule cutoffs, for 0.959 and 0.055. Every false alarm at 0.7 comes from the boolean-name rule (0.400 of its five clean cases); `threshold: 0.9` on that rule moves it to recall 0.750 and 0.000 — one catch traded for two false alarms, which is why the tier and the cutoff are per-rule choices, not new defaults. A one-sentence wording change (new text breaking the rule again is a violation even where the file already breaks it) was tried, made the same-file cases score high, but did not move the known miss and cost a tune catch on the shared set, so it was reverted. **Case `r15-04` (a new `var` in a file that already declares one) is still a known miss: 0.35 before, 0.39 after, both below the cutoff, and the new same-file cases swing between 0.51 and 0.96 across runs of the shipped question.** The shared 153 cases are unmoved (all: 0.912 and 0.028). Default behaviour is unchanged with no header lines and `softThreshold: 0`.
The bench grew to 161 labelled cases (17 rules: 107 tune, 54 holdout). A rule body may now start with `threshold:` and `severity:` lines beside `paths:`; a rule with its own cutoff fires there instead of `rules.threshold`, and severity orders findings high, normal, low, then score, in the steer and the trace. An opt-in `rules.softThreshold` (`0`, off) turns a score between it and a rule's cutoff into a soft "check whether this applies" sentence in the same steer, without holding and without counting as a finding. A full run at the shipped settings (161 cases, 147 asked: 74 violation, 73 clean) reaches recall 0.865, false alarms 0.027 and precision 0.970 at 0.7 (tp 64 / fp 2 / fn 10, the flat-0.7 view; the shipped per-rule thresholds give tp 63 / fp 0 / fn 11); the soft tier at 0.5 adds 8 catches and 4 false alarms against the per-rule cutoffs, for 0.959 and 0.055. Every false alarm at 0.7 comes from the boolean-name rule (0.400 of its five clean cases); `threshold: 0.9` on that rule moves it to recall 0.750 and 0.000 — one catch traded for two false alarms, which is why the tier and the cutoff are per-rule choices, not new defaults. A one-sentence wording change (new text breaking the rule again is a violation even where the file already breaks it) was tried, made the same-file cases score high, but did not move the known miss and cost a tune catch on the shared set, so it was reverted. **Case `r15-04` (a new `var` in a file that already declares one) is still a known miss: 0.35 before, 0.39 after, both below the cutoff, and the new same-file cases swing between 0.51 and 0.96 across runs of the shipped question.** The shared 153 cases are unmoved (all: 0.912 and 0.028). Default behaviour is unchanged with no header lines and `softThreshold: 0`.

### Turn question calibration (2026-09-27, first measurement)

Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "pi-warden",
"version": "0.74.1",
"version": "0.75.0",
"description": "Makes the Pi agent follow your project's rules. Jev judges every write against your pi-warden.md and quotes the broken rule back to the agent, names slop, breaks stuck loops, calls out unverified done claims, compresses large tool output, and holds the rare destructive command. Built on pi-typesafe.",
"type": "module",
"license": "MIT",
Expand Down
3 changes: 2 additions & 1 deletion src/config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -487,7 +487,8 @@ export function defaultConfig(): WardenConfig {
tools: [...COMMAND_TOOLS, "write", "edit"],
failOpen: true,
timeoutMs: 5000,
irreversible: { warn: 0.5, confirm: 0.7 },
// 0.9 holds: below it the judge is wrong one call in two to one in seven, and the 0.7 to 0.9 band held no call the user regretted.
irreversible: { warn: 0.5, confirm: 0.9 },
offTask: { warn: 0.6, steer: 0.85 },
intentMismatch: 0.9,
visibleMismatch: 0.8,
Expand Down
Loading
Loading