Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,14 @@ How to keep this current: add the entry in the same pull request as the change,

## Unreleased

## 0.89.0

### Changed

- Approval round 2 replaces the one-question design of 0.88.0. The approval request of a held call carries a second question, `reply_points_at_action` (do the agreeing parts of the reply point at this action, not at another item or question), and the call is released only when both `approved` and `reply_points_at_action` are at least 0.7. `replyApprovalQuestion` holds both questions; `askApproval` also returns `pointsAtAction`, and `Judgment.pointsAtAction` records it. No exported name changed.
- `asked` is the text of every assistant message after the previous user message and before the reply, in order, redacted, last 3,000 characters (before: the newest assistant message, last 1,500). A reply that arrives mid-run follows messages that hold only tool calls, and an explanation can sit earlier in the turn. The approval request now sends up to 3,000 redacted characters of the agent's turn; the consent text (`disclosure`) and `docs/data-handling.md` say so.
- Measured on 30 cases, 17 held-out cases, and 24 recorded holds (runs; `docs/guards.md`): wrong releases fall from 12 to 1 and correct releases from 81 to 78; held approvals rise from 9 to 12. The pre-set rule counted in runs chose the one-question design (3 fewer correct releases, 2 allowed); the owner shipped round 2 because the held-out cases show 1 wrong release against 6 at equal correct releases, and a wrong release costs more than a second "yes".

## 0.88.0

### Added
Expand Down
2 changes: 1 addition & 1 deletion docs/data-handling.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ With consent, requests go to `https://api.typesafe.ai` (default), or to the host

| Guard | Sent |
| --- | --- |
| **Action** | Your latest prompt (1500 characters), the task spine it is judged against (the thread's first request and up to four earlier requests, redacted, capped at 1200 characters together), up to eight earlier user and assistant messages (750 redacted characters each), the agent's text from the message that makes the call (500 redacted characters), the tool name, the command (2000 characters) or the file path (relative inside the project, `~`-shortened outside), whether the file exists, a 1500-character head/middle/tail sample of a `write` (or of the content a `bash` command writes to a file with the content in the command, with the written paths), the first three edit pairs (400 characters each) of an `edit`. Calls the ask gate decides offline (`action.ask`) send nothing at all, and so do read-only tools and read-only shell lines. The acting request carries none of the earlier user and assistant messages: no acting question reads them. The off-task, scope, and should-proceed questions are not on the acting request either, because nothing delivered reads their answers; one judged call in twenty (`action.traceSample`, default 0.05) sends a second request that carries them, with up to eight earlier user and assistant messages (750 redacted characters each) and the resolved rules content, which only that sampled request carries when no violation is open. `/warden test` sends one request and never the sample. The resolved active rules content (`pi-warden.md`, the configured files, or `AGENTS.md`/`CLAUDE.md`/`README.md` as fallback, token-aware truncated at ~4000 tokens) rides the acting request only while a violation is open on the call, because only the per-violation questions name it, and rides the sampled request above; with `rules.enabled` false no rules content leaves the machine. On the first guarded call after your reply, the tool names and commands (300 characters) or paths of up to six calls allowed in the previous turn, for the regret question. Only for a call that is held after an earlier hold, once your reply came in, one approval request with your reply, `asked` (the newest agent message before it, redacted, its last 1500 characters), the redacted call summary, and the reasons for the hold (pattern names and scores); the acting request carries no approval question. |
| **Action** | Your latest prompt (1500 characters), the task spine it is judged against (the thread's first request and up to four earlier requests, redacted, capped at 1200 characters together), up to eight earlier user and assistant messages (750 redacted characters each), the agent's text from the message that makes the call (500 redacted characters), the tool name, the command (2000 characters) or the file path (relative inside the project, `~`-shortened outside), whether the file exists, a 1500-character head/middle/tail sample of a `write` (or of the content a `bash` command writes to a file with the content in the command, with the written paths), the first three edit pairs (400 characters each) of an `edit`. Calls the ask gate decides offline (`action.ask`) send nothing at all, and so do read-only tools and read-only shell lines. The acting request carries none of the earlier user and assistant messages: no acting question reads them. The off-task, scope, and should-proceed questions are not on the acting request either, because nothing delivered reads their answers; one judged call in twenty (`action.traceSample`, default 0.05) sends a second request that carries them, with up to eight earlier user and assistant messages (750 redacted characters each) and the resolved rules content, which only that sampled request carries when no violation is open. `/warden test` sends one request and never the sample. The resolved active rules content (`pi-warden.md`, the configured files, or `AGENTS.md`/`CLAUDE.md`/`README.md` as fallback, token-aware truncated at ~4000 tokens) rides the acting request only while a violation is open on the call, because only the per-violation questions name it, and rides the sampled request above; with `rules.enabled` false no rules content leaves the machine. On the first guarded call after your reply, the tool names and commands (300 characters) or paths of up to six calls allowed in the previous turn, for the regret question. Only for a call that is held after an earlier hold, once your reply came in, one approval request with your reply, `asked` (the text of every agent message of the turn before your reply, in order, redacted, its last 3000 characters), the redacted call summary, and the reasons for the hold (pattern names and scores); the acting request carries no approval question. |
| **Rules** | The project-relative path, a 6000-character sample of a `write` (or of the content a `bash` heredoc, `echo`, or `printf` writes to a file, judged as a `write`) or each edit's new text (1500 characters) with about 40 lines of the current file around the replaced text, as they are and with the edit applied, and the rule text from your rules file or the condensed fallback document (`rules.maxChars`). No task text. Files under `rules.exclude` are never sent. `/warden rules check` sends each rule's id, heading, `paths:` scope, and text (redacted, clipped at 400 characters) in one request per 32 questions, with no file content and no task text. `/warden rules calibrate` sends one request per changed file of the sampled commits — the project-relative path, the added and removed lines with about 40 lines of the file after the commit around them (redacted), and the rule text — only after a confirm dialog that shows how many requests go out and the redacted diffs; a headless run sends nothing without the explicit `--yes`. `/warden rules tune` sends no request: it hands the flagged rules to the session's agent as one message. | `/warden rules audit` sends a 6000-character redacted sample of each selected file judged as a `write`, plus the rule text, and only after the confirm dialog or `--yes`; files under `rules.exclude` or `rules.skip` are never selected. `/warden bench` sends only one fixed built-in sample file and the rule text; no project content goes for it. |
| **Rules at turn start** | Before each new user message, when the rules guard and `rulesAtTurnStart` are on: the new request (1500 redacted characters), the task spine (the thread's first request and up to four earlier requests, redacted and capped as above), and the project's rule set — each rule's heading, its text (300 characters), and its `paths:` scope. One question per rule; no file content, no tool output, and no earlier assistant text. Nothing is sent when the prompt is a short continuation or a relayed child report (the two prompts the conscience's local gate also skips), when the rules source has no rule headings, or when judgments are off. When the request fails, times out, comes back after the run ended, or no rule passes the threshold, it has already been sent and nothing is appended to the session. |
| **Stuck** | The last 12 tool calls (300 characters each) with 400-character output tails, and, while `stuck.evidence` is on (the default), a structured `evidence` section: per run the parsed failing test, error, location, summary, exit code and which earlier run failed the same way (or 300 characters of head and 300 of tail when nothing parses), per `edit`/`write` the project path and a diff of the change capped at 600 characters, and a digest. Every string is redacted; the whole object is capped at 4 KB. |
Expand Down
2 changes: 1 addition & 1 deletion docs/extension-authors.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ What spans calls in a session lives in `ActionGuard` (hold, reply, retry approva

From 1.0, semver covers exactly these exports of the package root (`pi-warden`), together with the types of their parameters and results:

- Action guard: `evaluateAction`, `ActionGuard`, `describeAction`, `matchPatterns`, `isReadOnlyCommand`, `stripDataText`, `formatVerdict`, and the question sets `questions`, `intentQuestion`, `visibleQuestion`, `slopQuestions`, `approvalQuestion`, `securityQuestion`, `regretQuestions`, and the approval step for a held call: `settleApproval`, `askApproval`, `buildApprovalRequest`, `replyApprovalQuestion`, `describeAsked`, `APPROVAL_THRESHOLD`. `ActionGuard` and `evaluateAction` with `retryAfterHold` both call `settleApproval`; the field `asked` of `Conversation` and of `ActionInput` is the agent message the user's reply answers.
- Action guard: `evaluateAction`, `ActionGuard`, `describeAction`, `matchPatterns`, `isReadOnlyCommand`, `stripDataText`, `formatVerdict`, and the question sets `questions`, `intentQuestion`, `visibleQuestion`, `slopQuestions`, `approvalQuestion`, `securityQuestion`, `regretQuestions`, and the approval step for a held call: `settleApproval`, `askApproval`, `buildApprovalRequest`, `replyApprovalQuestion`, `describeAsked`, `APPROVAL_THRESHOLD`. `replyApprovalQuestion` holds both questions of the approval request, `approved` and `reply_points_at_action`; a call is released only when both reach `APPROVAL_THRESHOLD`, and `askApproval` returns both scores (`approved`, `pointsAtAction`; also `Judgment.pointsAtAction`). `ActionGuard` and `evaluateAction` with `retryAfterHold` both call `settleApproval`; the field `asked` of `Conversation` and of `ActionInput` is the agent's words the user's reply answers: the text of every assistant message of the turn before it, in order, last 3,000 characters.
- Rules: `evaluateRules`, `RuleStore`, `RulesGuard`, `parseRules`, `matchGlob`.
- Redaction: `redact`, `syntheticish`, `partitionSecrets`.
- Stuck detector: `AttemptWindow`, `makeAttempt`, `evaluateStuck`.
Expand Down
Loading
Loading