Skip to content

PLAN: evaluation rubric v2 — enforce couplings, verdict vocabulary, instrument fixes (from the three-way design review) #7

Description

@mmcky

Orientation, added 2026-08-07 — read this before picking anything up here. The product was reframed on 2026-08-07: triageshould this lecture be converted at all? — is now the skill's front door, and review mode is kept for when a conversion PR exists (#33, benchmark 0.4.0). Nothing on this issue was invalidated by that: #33 deliberately left the rubric, weights, gates, bands and scorecard JSON untouched, and both anchors reproduce byte-identically. But the five open items below were written when review mode was the product, so weigh them accordingly — the readability instrument in particular sits in the layer that got demoted, and the evidence for demoting it is that in every evaluation to date the recommendation was decided by the triage-layer instruments and never moved by the scorecard on top. This issue has not moved since 2026-07-27 and is the one genuinely untouched thread in the benchmark programme.

Two independent design critiques of the benchmark evaluation system — a fresh unframed review session and a 36-agent adversarial workflow in which every critique faced a steelman defense of the original design — are merged in reviews/benchmark-design-2026-07-21-merged.md (independent report alongside it). This issue tracks the v2 changes that survived the defense. The measurement foundation (as-used doctrine, evidence→engine determinism, verbatim extraction) was independently endorsed by both reviews and is not in scope — this is a scoring-superstructure revision.

Corrections of record already applied on the PR #6 branch (findings briefly posted to lecture-python.myst#654 were withdrawn — that PR gets one authoritative evaluation after v2): the markov_asset lecture does execute in notebook order — a stale global masks the stray err.throw(), silently disabling the checkify stability validation (subtler than the "does not build" we claimed); both reference replays deviate from their lectures' construction patterns; the as-used totals are single passes, not medians.

Engine changes (validatable against the committed evidence files)

  • Enforce the safety couplings in code — derive the logic-design bug cap from builds/x64-divergence inside score_all; verdict gates: correctness 1 caps at "net regression", correctness 2 caps at "mixed/wash". Closes every demonstrated hole incl. the honest-evidence 4.2 (float32 catastrophe, no logic bug → "clear improvement — merge")
  • "No-conversion" verdict output — reconciles review and triage: a polished ge_arrow currently scores 4.0 "merge" while triage says don't-convert at a 0.028 s baseline
  • Sensitivity stamp in score.py — perturb each boolean and band-adjacent value; mark scorecards robust/fragile with deciding flips listed (one concept currently flips markov_asset's band; one boolean flips ge_arrow's)
  • K-repeat as-used measurement — median of ≥3 fresh-process runs, contested-band annotation when the spread crosses an edge

Documentation honesty pass

  • Thresholds labeled as policy choices with derivation notes ("calibrated" → "anchored"; the wash band's noise-floor rationale recorded)
  • Fan-out paragraph: a root-cause fact's total influence is the sum over its manifestations — currently unstated and the weight table invites misreading; dedupe the near-verbatim vectorisation criterion shared by logic and style
  • Concept-grain rule stated explicitly (one item per reader-facing API surface element, symmetric old/new — currently exists only by example)

Design input needed from @xuanguang-li

  • Readability instrument v2 — the one critique that fully survived: docstring coverage inverts the framework's own exemplars (odu.py, the LOW example, measures 0.86; flagship-aiyagari style measures ~0.41–0.55 and caps at 2). Proposal: equation-traceability (per numbered equation, cite the single implementing expression; fraction traceable old vs new), with concept lists moving into evidence.json as cited slots (extends the M1 proposal on Land the lecture evaluation system (benchmark plugin 0.3.0: rubric v2, skill wired) #5)
  • Extraction/replay verification — a mechanical step diffing extracted code and replayed call sequences against the lecture's actual cells; longer-term, evaluate executing the lecture itself at both refs (nbclient per-cell timings, PROJECT: QuantEcon benchmarking programme (code performance & execution) meta#335 telemetry) as the as-used source

Explicitly defended — no change

The efficiency ratio form (goal-failure semantics; log-rescoring changes no committed verdict), the weighted total (the four-gate alternative was shown to be a lossy projection fitted on the rubric's own calibration set), min() aggregation's Goodhart-resistant shape, and per-consequence multi-counting as a principle.

Relates to #4 (skill wiring — the v2 rubric changes what the skill drives) and #5 (the landing PR carrying the corrections).

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementImprovement to existing content or functionality

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions