Steer calibration with blind labels: off-task and should-proceed stay trace-only - #151
Closed
DevMortimer wants to merge 7 commits into
Closed
DevMortimer wants to merge 7 commits into
DevMortimer wants to merge 7 commits into
Conversation
DevMortimer
marked this pull request as ready for review
September 30, 2026 14:07
DevMortimer
marked this pull request as draft
September 30, 2026 14:14
…ate term from a script comment
DevMortimer
marked this pull request as ready for review
September 30, 2026 23:00
DevMortimer
marked this pull request as draft
September 30, 2026 23:08
Owner
Author
|
Closing: the content of this pull request (its docs, README rows, and scripts, with the labeller corrected to one language model) is included in #162, without the version bump. |
DevMortimer
added a commit
that referenced
this pull request
Oct 1, 2026
#162) Makes the README and docs true for 1.0 and describes what 1.0 ships. Draft: no version bump, and nothing under `src/` changes (`git diff origin/main...HEAD -- src` is empty). The A/B results and the 1.0.0 bump come later. ## What changed - **Steer calibration (#151) brought in**: its docs, README rows, `scripts/steer-calibration.mjs`, `scripts/steer-ask-probe.mjs` and `should_ask`. Not its version bump or CHANGELOG version heading. Corrected: the labels were made by one language model, not "by hand", and a script comment no longer uses an internal term. - **"By hand" claims removed.** 140 intent-mismatch labels (README, guards.md, CHANGELOG): the record (#139) does not say who labelled them, so the text now says that. The 14-call sample in the old field report: same. - **`scripts/field-usage.mjs`**: the steer classifier knows eight more message kinds. "Other" steers, 2026-09-16 to 2026-10-01: 239 of 1,834 (13.0%) → 10 (0.5%). 2026-09-25 to 2026-10-01: 152 of 473 (32.1%) → 5 (1.1%). The JSON carries rule counts, not rule names. - **New field report** `eval/reports/2026-10-01-field-usage/` (aggregate counts only); hero image and Receipts refreshed from it (`scripts/render-hero.mjs`). - **README**: 1.0 features, a "What it costs" section, the Jev/offline-floor sentence fixed, off-task and should-proceed rows, Learning from holds matches #149. - **docs/guards.md**: Action guard step 3 describes the acting request, trace-only questions and ask gate as shipped. Calibration now records the ask gate and lean request, and three negative results: working-memory gate (checked against #152), relevance compaction hybrid (#150), stale-result stubs (#158). - **Doc mismatches fixed**: `docs/extension-authors.md` (`verdict.level` includes `"deny"`); `docs/data-handling.md` (conscience local ranking and what sends nothing; allowed rows keep judge data, held rows keep the summary); `docs/commands.md` (`/warden status` lifetime hold counts); `docs/configuration.md` (off-task and should-proceed rows). The other claims in those four docs matched `src/`. ## Numbers in the README and their sources | Number | Source | | --- | --- | | 1,168 sessions; 419 holds; 124 done-check nudges, 94 followed by a check (76%); 111 holds with an outcome: 79 safer route, 29 approved, 3 declined; 422 rule steers over 41 rules | `eval/reports/2026-10-01-field-usage/` (`scripts/field-usage.mjs`) | | 18,075 calls replayed; 0.27% held at 0.7, 0.1% at 0.9 | `eval/reports/2026-09-21-calibration-0.33.3/report.md` | | 150 paired runs, 6 vs 0 rule breaks | `runs.json` of the four `eval/reports/2026-09-18T00-*` batches: 45 + 30 + 45 + 30 = 150 runs per cell; control runs with a violation 1 + 3 + 0 + 2 = 6, warden 0 | | 13,952 cases over 109 cycles | `eval/reports/2026-09-18-overnight-stability/report.md` (line 4) | | 15,346 calls, 15% / under 1%, about 52 false alarms | docs/guards.md, Irreversible hold threshold | | AUROC 0.815, 140 calls, 36 of 37 steers | #139; docs/guards.md, Intent mismatch | | 48 compactions, 4.7 times, 1 of 34 files | docs/guards.md, Relevance compaction replay | | Ask gate: 28,036 → 14,326 requests (−48.9%); 64.2M → 20.5M tokens (−68.1%); 327 of 338 (96.7%) | #153; guards.md, Ask gate and lean request | | Conscience: 2 → 1 request per prompt; 14,620 → 10,241 tokens (−30%); 30.9% of 7,614 prompts send none | #155; guards.md, Local gate, top k, and tip text | | Prompt hold about 255 ms → under 1 ms (p90); reminder precision 68.9%, recall 21.5% | #152; guards.md, Turn-start delivery calibration and Rules at turn start calibration | | Approval questions per 1,000 judged calls 90.0 → 0.7; wrong releases 45 → 1; round 2 sets and totals | #160, #161; guards.md, Approval on demand | | Hold database 285.9 MiB → 41.6 MiB on 52,362 rows; 90-day prune of allowed rows, 365 for holds | #149; docs/data-handling.md; `learning.*` in docs/configuration.md | | Stale stubs: safe point in 10 of 1,042 sessions, median 0.00% | #158; guards.md, Stale-result stubs | | Hybrid compaction: 5 of 34 vs gate of 10, 1.28x; replace 5.02x, 11 of 34 | #150; guards.md, Relevance compaction hybrid | | Working memory: best cut drops 23.5%, misses 12.8% vs gate 30% / 10% | guards.md, Working-memory feasibility | | Steer calibration: 427 labelled calls, 2 needing a question, no slice passes | guards.md, Steer calibration | ## Still to do before ready for review - `npm run check` result with its test count (not run on this commit yet). #151 content is included here; it will be closed with a pointer. ## Second pass (docs checked against src) - `docs/configuration.md` vs `src/config.ts`: the defaults JSON and the widget templates match `defaultConfig()` key for key. Fixed: `compaction.maxRequests` default 12 (doc said 20); the project-file table now lists `action.ask.enabled` and `rulesAtTurnStart.enabled`/`.threshold` as stricter-only; user-only keys named (`typesafeBackend`, `context.filter.enabled`, `security.maskOutput`, `widget`, `steers`, `learning`, `conscience`, `steerVisible`, `notices`, `steerBudget`, rule lists, `action.floor`); added `stuck.diffLimit`, `stuck.tailLimit`, `context.compactAppendix`, `PI_WARDEN_DB`, `PI_WARDEN_STEER_STATS`, `PI_WARDEN_INDEX_DIR`; bounds for `runaway.repeats`, `waste.every`, `visibleMismatch`; index path typo; a table that lost its header. - `docs/guards.md` vs `src/`: caps and limits checked (250/750/1200/1500/2000/3000/6000 characters, 256/24 runaway, spine, steer kinds). Fixed: adaptive-steer kind list now matches `STEER_KINDS` and `NEVER_MUTED`. Calibration sections are measurements and were not re-run. - `npm run check`: typecheck, build, and 1412 tests pass (1412 pass, 0 fail). `main` was already merged. - Added `docs/upgrading.md` (0.74.1 to 1.0) with a README pointer. Each point was checked against `src/` and the 0.74.1 tag; the 1.0 config warning for the two removed keys is not in this PR, so the page says only that they are ignored. No version bump. The last commit bumps the version to 1.0.0, after merging the config warnings for the two removed settings; the upgrade guide names that warning.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
action.offTask.warn,action.offTask.steer,action.shouldProceed.threshold, andaction.shouldProceed.steerchange only the reason text and the trace counts (docs/guards.md,docs/configuration.md).scripts/steer-calibration.mjsdraws a blind sample from a read-only copy of the hold log and scores the labels against candidate steer slices.scripts/steer-ask-probe.mjsmeasures a candidate question on the same labels.should_askjoins the candidate questions inscripts/action-candidates.mjs: recorded only, never acted on.No guard behavior changed and no steer ships.
README.mdstill describes an off-task steer at 0.85; this change does not touch that file.Measurement
Window: 27,387 judged calls recorded since 2026-09-25; 8,697 of them (32%) at or below 0.6 on
should_proceed. 427 calls sampled in strata (150 mutating calls with scopeunrelated, stratified by score; 50plausible side step; 50 with no off-task reason; 150 at or below 0.6, stratified by score; 50 above) and labelled by hand without scores or strata in view. Off-task means "the call does not serve the user's request or a step it needs"; should-ask means "a careful engineer would ask the user before this call"; uncertain was allowed. Samples and labels stay outside the repository.Gate to ship a steer: precision at or above 0.80 on at least 40 labelled calls in the slice, and at most 5 steers per 1,000 judged calls.
No slice passes. The labels hold no off-task call in 427 (95% upper bound below 1%) and two should-ask calls: merging two pull requests that no instruction named, and reading a third-party API key from a live deployment for a side experiment. Both steers therefore stay trace-only, as the code already has them.
Blocked measurement
The proposed rewording (
should_ask) is meant to be measured on these same labels byscripts/steer-ask-probe.mjs, one request per call, next to the shipped question. The run cannot complete: the TypeSafe account answers HTTP 402 (no balance). Smallest unblocking action: fund the account, or export a fundedTYPESAFE_API_KEY, then run the probe with--yesand score it with--score. Until those numbers exist the question stays extra-only.Checks
npm run checkon this branch: typecheck clean, 1263 of 1263 tests pass, build clean.npm run checkon the base commit: typecheck clean, 1263 of 1263 tests pass, build clean.