Skip to content

Approval round 2: a second question and the whole turn as asked - #161

Merged
DevMortimer merged 12 commits into
mainfrom
carvel/deckhand-b/approval-context
Oct 1, 2026
Merged

DevMortimer merged 12 commits into
mainfrom
carvel/deckhand-b/approval-context

Conversation

@DevMortimer

Copy link
Copy Markdown
Owner

Switches approval on demand to round 2 (0.89.0). 0.88.0 shipped the one-question design.

What changes

  • The approval request of a held call asks two questions: approved and reply_points_at_action (do the agreeing parts of the reply point at this action, not at another item or question). The call is released only when both are at least 0.7.
  • asked is the text of every assistant message after the previous user message and before the reply, redacted, last 3,000 characters (was: the newest assistant message, last 1,500). A reply that arrives mid-run follows messages that hold only tool calls, and an explanation can sit earlier in the turn.
  • Both paths share the step: ActionGuard and evaluateAction with retryAfterHold call settleApproval. No exported name changed; replyApprovalQuestion holds the new question, askApproval also returns pointsAtAction, Judgment.pointsAtAction records it.
  • The consent text (disclosure) and docs/data-handling.md say the approval request sends up to 3,000 redacted characters of the agent's turn.

Measurement (runs = cases × 3, threshold 0.7; full tables in docs/guards.md)

Set Design Released right Released wrong Held right Held wrong
30 cases one question 39 3 48 0
round 2 37 0 51 2
17 held-out one question 21 6 24 0
round 2 21 1 29 0
24 recorded holds one question 21 3 39 9
round 2 20 0 42 10
Total one question 81 12 111 9
round 2 78 1 122 12

The rule and the decision. The rule, set before round 2 was measured: ship round 2 only if its wrong releases are fewer and its correct releases are at most 2 fewer. Counted in runs, round 2 has 3 fewer correct releases, so the rule chose the one-question design; counted by case the gap is 2. The rule did not name its unit. Round 2 ships anyway: on the held-out cases, committed before the measurement, it releases as many correct cases (21 against 21) with 1 wrong release against 6, and in total it removes 11 of 12 wrong releases and adds 3 holds that need a second "yes". A wrong release runs a held destructive call without consent; a held approval costs one more reply.

Check that the shipped code reproduces the numbers. The 24 recorded holds, one run each, through askedBeforeReply and settleApproval, counted by hold: released right 7, released wrong 0, held right 14, held wrong 3. The measurement gives 6/0/14/4, within one hold. 24 requests, 43,277 input and 936 output tokens.

Tests. Each of the three release combinations (both high, approved only, pointing only) on the Action guard and the library path; asked with several assistant messages, a newest message with only tool calls, clipping that keeps the end, and redaction. npm run check passes (1,367 tests).

@DevMortimer
DevMortimer marked this pull request as ready for review October 1, 2026 02:33
@DevMortimer
DevMortimer merged commit 5ff3c77 into main Oct 1, 2026
6 checks passed
DevMortimer added a commit that referenced this pull request Oct 1, 2026
#162)

Makes the README and docs true for 1.0 and describes what 1.0 ships.
Draft: no version bump, and nothing under `src/` changes (`git diff
origin/main...HEAD -- src` is empty). The A/B results and the 1.0.0 bump
come later.

## What changed

- **Steer calibration (#151) brought in**: its docs, README rows,
`scripts/steer-calibration.mjs`, `scripts/steer-ask-probe.mjs` and
`should_ask`. Not its version bump or CHANGELOG version heading.
Corrected: the labels were made by one language model, not "by hand",
and a script comment no longer uses an internal term.
- **"By hand" claims removed.** 140 intent-mismatch labels (README,
guards.md, CHANGELOG): the record (#139) does not say who labelled them,
so the text now says that. The 14-call sample in the old field report:
same.
- **`scripts/field-usage.mjs`**: the steer classifier knows eight more
message kinds. "Other" steers, 2026-09-16 to 2026-10-01: 239 of 1,834
(13.0%) → 10 (0.5%). 2026-09-25 to 2026-10-01: 152 of 473 (32.1%) → 5
(1.1%). The JSON carries rule counts, not rule names.
- **New field report** `eval/reports/2026-10-01-field-usage/` (aggregate
counts only); hero image and Receipts refreshed from it
(`scripts/render-hero.mjs`).
- **README**: 1.0 features, a "What it costs" section, the
Jev/offline-floor sentence fixed, off-task and should-proceed rows,
Learning from holds matches #149.
- **docs/guards.md**: Action guard step 3 describes the acting request,
trace-only questions and ask gate as shipped. Calibration now records
the ask gate and lean request, and three negative results:
working-memory gate (checked against #152), relevance compaction hybrid
(#150), stale-result stubs (#158).
- **Doc mismatches fixed**: `docs/extension-authors.md` (`verdict.level`
includes `"deny"`); `docs/data-handling.md` (conscience local ranking
and what sends nothing; allowed rows keep judge data, held rows keep the
summary); `docs/commands.md` (`/warden status` lifetime hold counts);
`docs/configuration.md` (off-task and should-proceed rows). The other
claims in those four docs matched `src/`.

## Numbers in the README and their sources

| Number | Source |
| --- | --- |
| 1,168 sessions; 419 holds; 124 done-check nudges, 94 followed by a
check (76%); 111 holds with an outcome: 79 safer route, 29 approved, 3
declined; 422 rule steers over 41 rules |
`eval/reports/2026-10-01-field-usage/` (`scripts/field-usage.mjs`) |
| 18,075 calls replayed; 0.27% held at 0.7, 0.1% at 0.9 |
`eval/reports/2026-09-21-calibration-0.33.3/report.md` |
| 150 paired runs, 6 vs 0 rule breaks | `runs.json` of the four
`eval/reports/2026-09-18T00-*` batches: 45 + 30 + 45 + 30 = 150 runs per
cell; control runs with a violation 1 + 3 + 0 + 2 = 6, warden 0 |
| 13,952 cases over 109 cycles |
`eval/reports/2026-09-18-overnight-stability/report.md` (line 4) |
| 15,346 calls, 15% / under 1%, about 52 false alarms | docs/guards.md,
Irreversible hold threshold |
| AUROC 0.815, 140 calls, 36 of 37 steers | #139; docs/guards.md, Intent
mismatch |
| 48 compactions, 4.7 times, 1 of 34 files | docs/guards.md, Relevance
compaction replay |
| Ask gate: 28,036 → 14,326 requests (−48.9%); 64.2M → 20.5M tokens
(−68.1%); 327 of 338 (96.7%) | #153; guards.md, Ask gate and lean
request |
| Conscience: 2 → 1 request per prompt; 14,620 → 10,241 tokens (−30%);
30.9% of 7,614 prompts send none | #155; guards.md, Local gate, top k,
and tip text |
| Prompt hold about 255 ms → under 1 ms (p90); reminder precision 68.9%,
recall 21.5% | #152; guards.md, Turn-start delivery calibration and
Rules at turn start calibration |
| Approval questions per 1,000 judged calls 90.0 → 0.7; wrong releases
45 → 1; round 2 sets and totals | #160, #161; guards.md, Approval on
demand |
| Hold database 285.9 MiB → 41.6 MiB on 52,362 rows; 90-day prune of
allowed rows, 365 for holds | #149; docs/data-handling.md; `learning.*`
in docs/configuration.md |
| Stale stubs: safe point in 10 of 1,042 sessions, median 0.00% | #158;
guards.md, Stale-result stubs |
| Hybrid compaction: 5 of 34 vs gate of 10, 1.28x; replace 5.02x, 11 of
34 | #150; guards.md, Relevance compaction hybrid |
| Working memory: best cut drops 23.5%, misses 12.8% vs gate 30% / 10% |
guards.md, Working-memory feasibility |
| Steer calibration: 427 labelled calls, 2 needing a question, no slice
passes | guards.md, Steer calibration |

## Still to do before ready for review

- `npm run check` result with its test count (not run on this commit
yet).


#151 content is included here; it will be closed with a pointer.



## Second pass (docs checked against src)

- `docs/configuration.md` vs `src/config.ts`: the defaults JSON and the
widget templates match `defaultConfig()` key for key. Fixed:
`compaction.maxRequests` default 12 (doc said 20); the project-file
table now lists `action.ask.enabled` and
`rulesAtTurnStart.enabled`/`.threshold` as stricter-only; user-only keys
named (`typesafeBackend`, `context.filter.enabled`,
`security.maskOutput`, `widget`, `steers`, `learning`, `conscience`,
`steerVisible`, `notices`, `steerBudget`, rule lists, `action.floor`);
added `stuck.diffLimit`, `stuck.tailLimit`, `context.compactAppendix`,
`PI_WARDEN_DB`, `PI_WARDEN_STEER_STATS`, `PI_WARDEN_INDEX_DIR`; bounds
for `runaway.repeats`, `waste.every`, `visibleMismatch`; index path
typo; a table that lost its header.
- `docs/guards.md` vs `src/`: caps and limits checked
(250/750/1200/1500/2000/3000/6000 characters, 256/24 runaway, spine,
steer kinds). Fixed: adaptive-steer kind list now matches `STEER_KINDS`
and `NEVER_MUTED`. Calibration sections are measurements and were not
re-run.
- `npm run check`: typecheck, build, and 1412 tests pass (1412 pass, 0
fail). `main` was already merged.


- Added `docs/upgrading.md` (0.74.1 to 1.0) with a README pointer. Each
point was checked against `src/` and the 0.74.1 tag; the 1.0 config
warning for the two removed keys is not in this PR, so the page says
only that they are ignored. No version bump.



The last commit bumps the version to 1.0.0, after merging the config
warnings for the two removed settings; the upgrade guide names that
warning.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant