Orientation, added 2026-08-07 — read this before picking anything up here. The product was reframed on 2026-08-07: triage — should this lecture be converted at all? — is now the skill's front door, and review mode is kept for when a conversion PR exists (#33, benchmark 0.4.0). Nothing on this issue was invalidated by that: #33 deliberately left the rubric, weights, gates, bands and scorecard JSON untouched, and both anchors reproduce byte-identically. But the five open items below were written when review mode was the product, so weigh them accordingly — the readability instrument in particular sits in the layer that got demoted, and the evidence for demoting it is that in every evaluation to date the recommendation was decided by the triage-layer instruments and never moved by the scorecard on top. This issue has not moved since 2026-07-27 and is the one genuinely untouched thread in the benchmark programme.
Two independent design critiques of the benchmark evaluation system — a fresh unframed review session and a 36-agent adversarial workflow in which every critique faced a steelman defense of the original design — are merged in reviews/benchmark-design-2026-07-21-merged.md (independent report alongside it). This issue tracks the v2 changes that survived the defense. The measurement foundation (as-used doctrine, evidence→engine determinism, verbatim extraction) was independently endorsed by both reviews and is not in scope — this is a scoring-superstructure revision.
Corrections of record already applied on the PR #6 branch (findings briefly posted to lecture-python.myst#654 were withdrawn — that PR gets one authoritative evaluation after v2): the markov_asset lecture does execute in notebook order — a stale global masks the stray err.throw(), silently disabling the checkify stability validation (subtler than the "does not build" we claimed); both reference replays deviate from their lectures' construction patterns; the as-used totals are single passes, not medians.
Engine changes (validatable against the committed evidence files)
Documentation honesty pass
Explicitly defended — no change
The efficiency ratio form (goal-failure semantics; log-rescoring changes no committed verdict), the weighted total (the four-gate alternative was shown to be a lossy projection fitted on the rubric's own calibration set), min() aggregation's Goodhart-resistant shape, and per-consequence multi-counting as a principle.
Relates to #4 (skill wiring — the v2 rubric changes what the skill drives) and #5 (the landing PR carrying the corrections).
🤖 Generated with Claude Code
Two independent design critiques of the benchmark evaluation system — a fresh unframed review session and a 36-agent adversarial workflow in which every critique faced a steelman defense of the original design — are merged in reviews/benchmark-design-2026-07-21-merged.md (independent report alongside it). This issue tracks the v2 changes that survived the defense. The measurement foundation (as-used doctrine, evidence→engine determinism, verbatim extraction) was independently endorsed by both reviews and is not in scope — this is a scoring-superstructure revision.
Corrections of record already applied on the PR #6 branch (findings briefly posted to lecture-python.myst#654 were withdrawn — that PR gets one authoritative evaluation after v2): the markov_asset lecture does execute in notebook order — a stale global masks the stray
err.throw(), silently disabling the checkify stability validation (subtler than the "does not build" we claimed); both reference replays deviate from their lectures' construction patterns; the as-used totals are single passes, not medians.Engine changes (validatable against the committed evidence files)
builds/x64-divergence insidescore_all; verdict gates: correctness 1 caps at "net regression", correctness 2 caps at "mixed/wash". Closes every demonstrated hole incl. the honest-evidence 4.2 (float32 catastrophe, no logic bug → "clear improvement — merge")Documentation honesty pass
Design input needed from @xuanguang-li
Explicitly defended — no change
The efficiency ratio form (goal-failure semantics; log-rescoring changes no committed verdict), the weighted total (the four-gate alternative was shown to be a lossy projection fitted on the rubric's own calibration set), min() aggregation's Goodhart-resistant shape, and per-consequence multi-counting as a principle.
Relates to #4 (skill wiring — the v2 rubric changes what the skill drives) and #5 (the landing PR carrying the corrections).
🤖 Generated with Claude Code