Skip to content

Add offline RQ2 to RQ4 analysis: OOS, uncertainty, LLM fallback, oracle - #17

Merged
drewOrc merged 5 commits into
mainfrom
feat/analysis
Sep 29, 2026
Merged

drewOrc merged 5 commits into
mainfrom
feat/analysis

Conversation

@drewOrc

@drewOrc drewOrc commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Step 4, second half (docs/PLAN.md sections 3 and 4, AC3, AC4, AC6): RQ2 to RQ4 from the stored logits and Haiku predictions. Nothing is trained and no API is called. No figures here; README and figures are step 5.

What it does

make analysis reads the 75 logits archives (BERT 18, of which k=100 are the three AC2 runs; ModernBERT 18; OOS ablation 3; baselines 36) and results/llm/haiku-8way.jsonl, and writes results/analysis/summary.json, curves.json and haiku.json. Two runs give byte-identical files: there is no randomness anywhere.

Before anything is computed: each curve index is re-verified with completeness.verify_index; the archive count (75) and the Haiku row count (8,600) are literals; within each split every archive holds the same gold labels, and Haiku's gold_intent equals them row by row. After: each output is written to a temporary file, read back and checked (25 groups, 3 seeds each, both Haiku splits) before it is moved into place, and only then is completed analysis (75/75 archives, 8600/8600 llm rows, 25 groups) printed.

Chosen on validation only (AC4)

Every choosing function raises LeakageError on any other split, and a test shuffles the test labels and checks that no choice moves.

  • Temperature: fit_temperature on validation. Not fitted for majority (constant scores).
  • 8-way aggregation: higher validation 8-way accuracy, ties go to argmax (the PLAN default). Both aggregations' test numbers are reported. Summed takes the argmax of temperature-scaled agent probabilities, since that argmax depends on T.
  • Deferral threshold per signal and target risk (2% and 5%): among the distinct validation scores, the largest coverage whose one-sided 95% Wilson upper bound (z = 1.645) on selective risk is at most the target. Tied scores are never split. No feasible threshold means every query goes to the LLM, recorded as feasible: false.
  • Hybrid signal: largest validation coverage at the target, then lower validation AURC.
  • Operating curve (high-confidence OOS misroute against tau): tau is the validation score at each coverage of a 5% grid.

Definitions

  • Signals (higher is more confident): msp and entropy (negative entropy) at T = 1, msp_t at the fitted T, margin the gap of the top two log-probabilities at the fitted T (for argmax that is (z1 - z2)/T, the same ranking as the raw score margin). TF-IDF gets msp_t and margin only and no uncalibrated ECE; majority gets none (PLAN section 4.1).
  • OOS misroute rate = 1 - OOS recall. High-confidence OOS misroute rate = gold OOS kept by the small model (confidence >= tau) and routed to an in-scope agent / gold OOS.
  • Oracle: defers exactly the queries the small model gets wrong. Error recovery rate = small-model errors deferred and fixed by Haiku / small-model errors. Recoverable caught = the same numerator / errors Haiku would fix.
  • AURC and the risk-coverage curve use the expected value under random order within tied scores; AUROC counts ties as one half; average precision is sklearn's definition (tests compare against sklearn).
  • Mean and sample std (ddof=1) over seeds 42, 43, 44.
  • Haiku: the new parser and the old project's substring rule (first hit, now in fixed AGENTS order; replies with two or more labels are counted as order-dependent), parse failures, and accuracy with parse failures counted wrong.

Main test numbers (8-way, %)

acc OOS recall LLM calls
Haiku alone 82.1 56.8 (Wilson 53.7 to 59.8) 100
ModernBERT k=100, small only 91.9 ± 0.1 61.1 ± 0.4 0
ModernBERT k=100, hybrid, target 2% 92.1 ± 0.1 62.4 ± 1.3 1.3 ± 1.6
ModernBERT k=100, oracle 95.1 ± 0.1 8.1 ± 0.1
ModernBERT k=10, hybrid, target 2% 88.0 ± 0.4 53.5 ± 1.5 23.9 ± 3.8

Haiku: 0 parse failures, 0 rows where the two parsers disagree.

Thresholds chosen on validation miss the target on test. At a 2% target, ModernBERT k=100 has a test selective risk of 7.34 ± 0.86 against 1.50 ± 0.13 on validation. The main result stays validation-only and is reported as is (option a); see the review section for the diagnosis.

Review fixes (round 1)

The review found the computation correct and asked for these, all done on this branch:

  • Guards no longer trust a split name alone (F2). select_threshold and coverage_thresholds take a Scored, built by Routed.scored(signal), so the name they check and the scores they use come from one routed split. A new test replaces the test logits with noise and checks that nothing chosen on validation moves: T, aggregation, every tau, the hybrid's signal, the operating-curve taus and the (b) taus.
  • RQ2 OOS detection uses its own score (F3). It is 1 - max in-scope probability at the fitted T (151 intents for argmax, 8 agents for summed). The four confidence signals stay for RQ3. summary.json has a generated ablation_comparison: detection with one fixed aggregation (argmax) for both models, router behavior from the final router.
  • Threshold diagnostics (F4). diagnostics.py covers the final hybrid of every encoder point. It reports validation and test selective risk, in-scope risk, the error rate of kept OOS rows, the test risk reweighted to the validation OOS share, and how much of the gap that share explains. It also reports option (b) as a sensitivity: validation reweighted to an 18.2% OOS share, a share known only from test, so diagnosis only.
  • F1 caller tests for the OOS AUROC and AUPRC direction (against sklearn). F5 hand-built choices for select_signal. F6 entropy direction. F7 a case where T changes the summed argmax. F8 a test row exactly at tau is kept (deferred_below). F9 std is null (not 0) when fewer than two seeds have a value, with n kept. F10 two indexes pointing at one archive fail. F11 no assert isinstance in production code.

Ablation, corrected (test, %, generated in ablation_comparison)

ModernBERT k=100 OOS AUROC OOS AUPRC hybrid OOS recall (2%) high-conf OOS misroute (2%) LLM calls (2%)
250 OOS training rows 98.33 ± 0.07 94.34 ± 0.25 62.43 ± 1.33 34.53 ± 5.53 1.27 ± 1.59
0 OOS training rows 97.80 ± 0.04 91.52 ± 0.26 40.60 ± 1.04 37.67 ± 2.90 12.07 ± 0.73

The earlier "higher AUROC without OOS training" came from ranking OOS by low confidence, which counts a confident oos prediction as the least OOS-like row.

Why thresholds miss on test (target 2%, test, %, from diagnostics)

point val risk test risk test reweighted to val OOS share gap explained kept OOS error val → test (b) test risk (b) coverage
ModernBERT k=10 1.54 ± 0.01 5.09 ± 1.12 2.11 ± 0.39 85.2 40.39 → 60.61 1.64 ± 0.21 61.01 ± 0.49
ModernBERT k=100 1.50 ± 0.13 7.34 ± 0.86 2.29 ± 0.11 86.7 19.91 → 36.48 2.60 ± 0.48 87.33 ± 1.85
BERT k=100 1.57 ± 0.00 5.98 ± 0.10 1.77 ± 0.08 95.3 24.81 → 35.48 2.15 ± 0.87 84.16 ± 5.03

Validation is 3.2% OOS and test is 18.2%. Most of the gap comes from that share; the rest is test OOS being harder to catch. Reweighting to the deployment OOS share still does not fully close it, so thresholds should be recalibrated on data close to real traffic before deployment.

After the refactor, curves.json is byte-identical to the previous commit. summary.json only gains fields, plus 48 std values that change from 0 to null (F9). Two runs of make analysis give identical files.

Review fixes (round 2)

  • R1: new test where coverage ties and the signal with the better AURC (msp_t) is listed after the worse one (msp) in SIGNALS. The old test's better signal also came first, so removing the AURC tie-break passed.
  • R2: new summed-aggregation test. When the oos agent carries the mass, oos_score is close to 1 and equals 1 minus the largest in-scope agent probability.
  • R3: the weighted branch of select_threshold (used only by the (b) sensitivity) now uses Kish's effective sample size. At each cut, p is the weighted error rate and n_eff = (sum w)^2 / sum w^2 over the kept rows; the bound uses k = p * n_eff. With unit weights this gives the unweighted result (tested), and with unequal weights the bound is wider than the nominal weighted count (tested). ModernBERT k=100 (b) at 2%: test risk 2.60 ± 0.48 (2.17, 3.12, 2.50), coverage 87.33 ± 1.85. It was 2.93 ± 0.63 and 88.55 ± 1.76.
  • R4: when seeds chose different signals, tau keeps the per-seed values and signal names instead of a mean and std. This covers the final hybrid, each aggregation's hybrid, the diagnostics and (b). Taus under by_signal share one signal and are still averaged.

Field-by-field comparison of summary.json against the previous commit: the only changes are the (b) fields (diagnostics.*.sensitivity_reweighted_validation) and how tau is written (R4). Every other number is unchanged, and curves.json is byte-identical.

Mutation check for this round (commit first, back up, mutate, restore from the backup). The control passes, and a known mutant fails. All of these fail: removing the AURC tie-break (R1); a summed score that keeps the oos agent (R2); the nominal count in place of Kish, two variants (R3); averaging taus across different signals, and by_signal inheriting the hybrid's signal (R4).

Tests

98 tests for the analysis modules (524 in total, all offline). They cover every metric against hand-computed values, the validation-only guards (shuffled test labels, noise test logits, split-name bypass), gold cross-checks, completion checks, and the diagnostics by hand (prior weights, kept rates, reweighting, gap explained, the (b) choice).

Mutation check (commit first, back up, mutate, restore from the backup):

  • Control: the harness passes with no change and fails on a known mutant.
  • Review survivors: all 9 fail now, as 10 mutants for F1, F2, F5 to F10 (F1 and F2 have two each).
  • New code: 14 mutants for F3, F4 and the Scored shape check. One survived the first pass (the diagnostics' high-confidence misroute ignoring whether the row was kept); a hand case now covers it and it fails.
  • Earlier set: 16 mutants still fail.

make lint, make test, make smoke, make verify-logits, make verify-llm and make analysis pass locally.

make analysis reads the 75 logits archives and the 8,600 stored Haiku
predictions, trains nothing and calls no API, and writes
results/analysis/{summary,curves,haiku}.json.

Temperature, 8-way aggregation, deferral threshold, hybrid signal and
operating-curve thresholds are chosen on validation only; every choosing
function raises LeakageError on any other split. Inputs are checked
before anything runs (curve indexes, archive count, Haiku rows, gold
labels equal across archives and Haiku), and outputs are read back and
checked before the completion line is printed.
Two mutants survived the first mutation pass: a two-sided z for the
threshold bound, and a temperature fit that skips the split check while
a later guard still catches test data. Each now has a direct test.
…ostics

Threshold choosers take a Scored built from one routed split, so a split
name and another split's arrays cannot be mixed. RQ2 OOS detection uses
1 - max in-scope probability at T; the ablation compares both models with
that score and one aggregation. A diagnostics block explains why
validation thresholds miss the target on test (OOS share, harder test
OOS) and reports the reweighted-validation sensitivity as diagnosis only.
Tests cover OOS score direction, noise test logits, signal choice, entropy
direction, temperature-scaled summed argmax, ties at tau, single-seed std,
and a shared archive across indexes.
The diagnostics hand case had no OOS row routed to an agent and deferred,
so dropping the kept condition went unnoticed.
… and summed OOS score

The weighted threshold (used only by the prior-shift sensitivity) now
bounds the weighted error rate with n_eff = (sum w)^2 / sum w^2 of the
kept rows instead of the nominal weighted count. Taus of seeds that chose
different signals keep per-seed values and signal names instead of a
mean. New tests: equal coverage goes to the lower AURC even when that
signal is listed later, and the summed OOS score is high when the oos
agent carries the mass.
@drewOrc
drewOrc merged commit 606837f into main Sep 29, 2026
4 of 5 checks passed
@drewOrc
drewOrc deleted the feat/analysis branch September 29, 2026 04:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant