Motivation
Where scores sit near the metric's upper bound, differences between strategies have little room to appear and a no-difference result carries weak information. The harness currently reports means, which hide saturation. It should detect the condition and offer a tier where it does not apply.
Proposal
- Report the score distribution per cell rather than the mean alone: share of runs at or above 0.95, and the empirical maximum.
- Flag cells where saturation makes a difference undetectable.
- Add a high-complexity tier — more entities, more relations, denser cross-references — sized so the strongest model lands clearly below the ceiling.
Done when:
Scope
- In scope: distributional reporting, saturation flagging, one additional complexity tier, baseline verification run.
- Out of scope: replacing the existing tiers; a full factorial re-run on the new tier — a baseline suffices to confirm headroom.
Open questions
- Is 0.95 the right saturation threshold, or should it derive from the observed ground-truth noise floor? If ground truth itself is imperfect, the effective ceiling is below 1.0.
- Should complexity increase via source-text length, entity and relation count, or structural density such as cross-references and cycles? These stress different failure modes.
- Does the new tier need its own ground truth built the same way as existing tiers, given the thesis-era provenance differences?
Related
Motivation
Where scores sit near the metric's upper bound, differences between strategies have little room to appear and a no-difference result carries weak information. The harness currently reports means, which hide saturation. It should detect the condition and offer a tier where it does not apply.
Proposal
Done when:
Scope
Open questions
Related