Skip to content

DeferLab: a world model is scored on when it declines to predict - #3

Open
di-omics wants to merge 5 commits into
feat/kinematics-and-stationsfrom
feat/world-model-screening
Open

di-omics wants to merge 5 commits into
feat/kinematics-and-stationsfrom
feat/world-model-screening

Conversation

@di-omics

@di-omics di-omics commented Aug 4, 2026

Copy link
Copy Markdown
Owner

Stacked on #2, so the diff here is only Phase 4 and the benchmark.

What this claims

Every world-model benchmark scores a rollout that was produced. This one scores the decision not to produce one: a correct refusal on an episode nothing could predict earns exactly what a correct prediction earns, no score is published for a model that ignores its action argument, and the same frozen threshold is what an arm path would gate on.

That gap is real as of Aug 2026, not assumed. A survey of roughly fifty world-model benchmarks (arXiv:2606.15032) names uncertainty calibration as one of seven required decision axes and finds it in one row. The closest existing work, RoboTrustBench (2606.01600), carries one refusal criterion of thirteen and, when a model actually refused to generate, excluded those samples as missing data rather than scoring them as correct abstentions.

What is in it

  • plr_lr/learn/ -- Phase 4: record, model, plan, evaluate. CPU-only, no weights, no arm.
  • plr_lr/learn/predict.py -- the contract a world model implements, carrying frames with hashes and capture times, plus the interventional gate that refuses rather than low-scores a model ignoring its actions.
  • plr_lr/learn/screen.py -- admit or defer, with deferral terminal and no fallback action.
  • plr_lr/learn/selective.py -- stdlib-only scoring core. scripts/score_standalone.py is a single-file copy that runs on a bare python:3.9-slim with zero installs, verified in CI.
  • plr_lr/learn/bench/ -- DeferLab: 5 tasks, 10 null controls, 16 refusal rules, no cross-task aggregate.

Four measured defects fixed on the way in

before after
DynamicsFit.improved on pure noise True 2 of 40 None with SPLIT-UNSOUND
ReachTask 5 trials at default jitter 5 identical trials, rate 1.00, valid=True jitter 25.0, identical trials is a hard term in valid
SamplingPlanner at n=256 193 distinct, reported 256 reports n_distinct
KinestheticSession on an injected trajectory source=kinesthetic pose_source records which it was

Two traps this got wrong first and had to measure out

Unknowable labels must be constructed, not asserted. The first version asserted them, and at S1 horizon 1 all 80 of 80 episodes labelled unknowable were fully determined, because the hidden latent only acted from step 1. Also 50-79% of S5 and 9-36% of S3. The abstention metrics were measuring the labeller. Every unknowable episode is now advanced under the full hidden grid at its scored horizon and must separate pairwise by more than two tolerances, or the pack build raises and emits nothing.

A must-fail control never observed to fail is not a control. N8 was differenced against a copy of itself and printed exactly +0.0000 on every task, on any pack however leaky. It now proves its probe works on the half it memorised before claiming anything about the half it did not, and it fails on a deliberately leaky pack.

The first scorecard is bad, deliberately

The built-in analytic floor prints NO INFORMATION on 14 of 16 rows, is rankable on 2, and scores RefusalRecall 0.0000 almost everywhere: it never abstains on an unknowable episode and has no channel that could. Publishing that under the harness's own rules is the point. A good-looking first row would have been the warning sign.

Honest ceiling

A world model predicts the visible half forward and is structurally blind to what the reagent actually is. Nothing here belongs in a servo loop: V-JEPA 2-AC is 16 s per action on a 4090 and NVIDIA's own world-action-model number is 590-800 ms per chunk, against a 20 ms control period. Every row reads basis: simulated; no arm has been run.

Verification

1065 tests pass, ruff clean, house-rules clean over the full history, all four examples exit 0, and the standalone scorer produces a full scorecard on stock Python 3.9.6 with zero installs.

ASCII-only, no banned typography, no attribution trailers, and one author
across the whole history. All four are asserted over git ls-files and over
git log --all, not over the working tree, because the working tree is not
what a reader clones.

The check exists because its absence is expensive in exactly one direction.
Catching a stray character before a commit costs nothing; catching it after
means rewriting published history, and a rewrite is itself an event that has
to be explained. scripts/check_house_rules.py is stdlib-only so it runs
anywhere, and tests/test_house_rules.py seeds a violation of each of the
three content categories and asserts the checker catches it, because a
checker nobody has watched fail is a checker with no evidence behind it.

Also fixes the seven pre-existing lint findings this rule set surfaces, and
pins the ruff selection in pyproject so CI and a local run enforce the same
thing. The larger ruff families rewrite annotations in ways that conflict
with this project's Python 3.9 floor, so the set stays explicit rather than
inherited.
record -> model -> plan -> evaluate, all of it CPU-only with no weights and
no arm. The dynamics fit is linear least squares on [pose, delta, 1], small
enough to run in milliseconds and large enough to catch the systematic errors
that matter on a bench: a consistent undershoot, an axis that lags, a
delta-proportional loss from backlash.

Four places where the permissive answer is a confident wrong number, each
found by measurement rather than by reading:

DynamicsFit.improved returned True on pure noise 2 times in 40 under a
transition-level split, because consecutive transitions from one episode are
not independent and a held-out transition is a neighbour of a trained one.
It now consults split_by and returns None, not False, when the split is
unsound: the claim is unevaluable, which is a different answer from refuted.

ReachTask.start_jitter_mm defaulted to 0.0, so five trials were five copies
of one trial and reported success_rate 1.00 with valid=True. The identical
-trials note was text and did not reach the verdict. Jitter now defaults to
25.0 and identical trials is a hard term in valid.

SamplingPlanner drew 256 candidates and overwrote 63 of them with the same
aimed trajectory, then reported 256. It now reports n_distinct beside
n_samples. The sampling is unchanged; the number was the bug.

KinestheticSession stamped source=kinesthetic on a fully synthetic injected
trajectory, so a dataset claimed a human dragged the arm when nobody had.
pose_source now records which it was and refuses the claim it cannot support.

The learn tests take a backend that reports degrees and lives outside the
simulation hierarchy from tests/_doubles.py rather than from any vendor
driver. A learning package whose tests need a specific robot to run is
coupled to that robot, and this one is not.
Three pieces that turn a prediction into a decision about whether to move.

predict.py is the contract a world model implements. It carries frames with
their hashes and capture times, because a request with no slot for a frame
silently conditions every video prediction on whatever image a closure
happened to capture, and a staleness check on a timestamp nobody recorded
cannot fire. abstain is true if and only if the state is all-NaN and the
confidence is NaN, checked row for row: abstaining while banking a prediction
is the obvious way to game an abstention metric, so it is illegal rather than
discouraged. A model that emits actions is routed to the policy seam by a
constructor that raises, because a taxonomy written in prose is not a rule.

The interventional gate runs before any score is emitted. It perturbs a
complete basis of the perturbable columns plus a seeded random direction and
takes the largest response, because "reads its action argument" is an
existential claim: an isotropic probe drawn near-orthogonal to the one axis a
task answers on refuses a model that is working fine. Failure is a refusal,
not a low score. A low score still gets ranked, and a ranked
non-world-model is the exact mismatch this whole thing exists to indict.

selective.py is stdlib-only on purpose, and scripts/score_standalone.py is a
single-file copy of it. Somebody with a result file should be able to score it
without installing this package, and CI runs that copy on a bare
python:3.9-slim with no pip install line to keep the claim true. E-AURC is the
ranked metric and its oracle and random bands print in the same row, so the
band a score sits in is never a second lookup. No rate is printed below n=30
and none is rankable below n=200, with no override.

screen.py admits a prediction only if the model did not abstain, the
confidence clears a threshold frozen on a disjoint calibration split, the
horizon is inside the measured trust horizon, the readout succeeded and the
frame is recent enough. Deferral is terminal. There is no fallback action,
because converting uncertainty into a different action is the conversion the
screen exists to prevent.

A clamped action is not the action that was screened, so the ACT verdict does
not transfer to it and the step is recorded as ABSTAIN(CLAMPED). This is not
hypothetical: the existing random-policy null has 47 percent of its steps
clamped, and 1 - pi/6 = 0.476 of a cube lies outside its inscribed ball, so
the rate is arithmetic rather than accident. Without the downgrade, "low
confidence predicts failure" and "large deltas get clamped" are the same
measurement wearing different labels.
Every world-model benchmark scores a rollout that was produced. This one
scores the decision not to produce one. A correct refusal on an episode
nothing could predict earns exactly what a correct prediction earns, and
abstention credit is never printed without the false-alarm rate beside it,
because a one-sided abstention number is unfalsifiable.

Five tasks, ten null controls, sixteen refusal rules, no cross-task
aggregate. The pixel column exists under a header reading NOT RANKED for one
reason: so a control that replays ground-truth frames can top it outright and
still lose the ranking.

Two things a benchmark of this shape gets wrong by default, both of which
this one got wrong first and had to be measured out:

Unknowable episodes must be constructed, never asserted. The first version
asserted them, and at S1 horizon 1 all 80 of 80 episodes labelled unknowable
had a terminal state that was byte-identical for every value of the hidden
latent, because that latent only acted from step 1. Fifty to seventy-nine
percent of S5 and nine to thirty-six percent of S3 were the same. The
abstention metrics were measuring the labeller. Every unknowable episode is
now advanced under the full hidden grid at the horizon it is scored at, and
must separate by more than two tolerances pairwise and one from nominal, or
the pack build raises and emits nothing. Oracle success is evidence of
knowability; oracle failure is not evidence of unknowability, and only the
conclusive direction is read.

A must-fail control that has never been observed to fail is not a control.
N8 was differenced against a copy of itself and printed exactly +0.0000 on
every task, forever, on any pack however leaky. It now demonstrates its probe
works on the half it memorised before claiming anything about the half it did
not, and it fails on a deliberately leaky pack, which is the only evidence
that it can.

The first scorecard is bad and that is the point. The built-in analytic floor
prints NO INFORMATION on fourteen of sixteen rows, is rankable on two, and
scores RefusalRecall 0.0000 almost everywhere: it never abstains on an
episode nothing could predict, and it has no channel that could. Publishing
that under the harness's own rules is the credibility event. A good-looking
first row would have been the warning sign.

The honest ceiling is in the README rather than implied. A world model
predicts the visible half forward and is structurally blind to what the
reagent actually is. Nothing here runs usefully CPU-only except the harness,
nothing here belongs in a servo loop, and every row reads basis: simulated
until a hardware run record exists.
The trigger was push on [main] plus pull_request, which reads as full
coverage and is not. A branch whose pull request targets another branch can
be reviewed and merged with this workflow never having run on it.

Found by watching it happen: the first push of this branch ran house-rules
and nothing else, and a silent job is harder to notice than a red one.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant