Conversation
ASCII-only, no banned typography, no attribution trailers, and one author across the whole history. All four are asserted over git ls-files and over git log --all, not over the working tree, because the working tree is not what a reader clones. The check exists because its absence is expensive in exactly one direction. Catching a stray character before a commit costs nothing; catching it after means rewriting published history, and a rewrite is itself an event that has to be explained. scripts/check_house_rules.py is stdlib-only so it runs anywhere, and tests/test_house_rules.py seeds a violation of each of the three content categories and asserts the checker catches it, because a checker nobody has watched fail is a checker with no evidence behind it. Also fixes the seven pre-existing lint findings this rule set surfaces, and pins the ruff selection in pyproject so CI and a local run enforce the same thing. The larger ruff families rewrite annotations in ways that conflict with this project's Python 3.9 floor, so the set stays explicit rather than inherited.
record -> model -> plan -> evaluate, all of it CPU-only with no weights and no arm. The dynamics fit is linear least squares on [pose, delta, 1], small enough to run in milliseconds and large enough to catch the systematic errors that matter on a bench: a consistent undershoot, an axis that lags, a delta-proportional loss from backlash. Four places where the permissive answer is a confident wrong number, each found by measurement rather than by reading: DynamicsFit.improved returned True on pure noise 2 times in 40 under a transition-level split, because consecutive transitions from one episode are not independent and a held-out transition is a neighbour of a trained one. It now consults split_by and returns None, not False, when the split is unsound: the claim is unevaluable, which is a different answer from refuted. ReachTask.start_jitter_mm defaulted to 0.0, so five trials were five copies of one trial and reported success_rate 1.00 with valid=True. The identical -trials note was text and did not reach the verdict. Jitter now defaults to 25.0 and identical trials is a hard term in valid. SamplingPlanner drew 256 candidates and overwrote 63 of them with the same aimed trajectory, then reported 256. It now reports n_distinct beside n_samples. The sampling is unchanged; the number was the bug. KinestheticSession stamped source=kinesthetic on a fully synthetic injected trajectory, so a dataset claimed a human dragged the arm when nobody had. pose_source now records which it was and refuses the claim it cannot support. The learn tests take a backend that reports degrees and lives outside the simulation hierarchy from tests/_doubles.py rather than from any vendor driver. A learning package whose tests need a specific robot to run is coupled to that robot, and this one is not.
Three pieces that turn a prediction into a decision about whether to move. predict.py is the contract a world model implements. It carries frames with their hashes and capture times, because a request with no slot for a frame silently conditions every video prediction on whatever image a closure happened to capture, and a staleness check on a timestamp nobody recorded cannot fire. abstain is true if and only if the state is all-NaN and the confidence is NaN, checked row for row: abstaining while banking a prediction is the obvious way to game an abstention metric, so it is illegal rather than discouraged. A model that emits actions is routed to the policy seam by a constructor that raises, because a taxonomy written in prose is not a rule. The interventional gate runs before any score is emitted. It perturbs a complete basis of the perturbable columns plus a seeded random direction and takes the largest response, because "reads its action argument" is an existential claim: an isotropic probe drawn near-orthogonal to the one axis a task answers on refuses a model that is working fine. Failure is a refusal, not a low score. A low score still gets ranked, and a ranked non-world-model is the exact mismatch this whole thing exists to indict. selective.py is stdlib-only on purpose, and scripts/score_standalone.py is a single-file copy of it. Somebody with a result file should be able to score it without installing this package, and CI runs that copy on a bare python:3.9-slim with no pip install line to keep the claim true. E-AURC is the ranked metric and its oracle and random bands print in the same row, so the band a score sits in is never a second lookup. No rate is printed below n=30 and none is rankable below n=200, with no override. screen.py admits a prediction only if the model did not abstain, the confidence clears a threshold frozen on a disjoint calibration split, the horizon is inside the measured trust horizon, the readout succeeded and the frame is recent enough. Deferral is terminal. There is no fallback action, because converting uncertainty into a different action is the conversion the screen exists to prevent. A clamped action is not the action that was screened, so the ACT verdict does not transfer to it and the step is recorded as ABSTAIN(CLAMPED). This is not hypothetical: the existing random-policy null has 47 percent of its steps clamped, and 1 - pi/6 = 0.476 of a cube lies outside its inscribed ball, so the rate is arithmetic rather than accident. Without the downgrade, "low confidence predicts failure" and "large deltas get clamped" are the same measurement wearing different labels.
Every world-model benchmark scores a rollout that was produced. This one scores the decision not to produce one. A correct refusal on an episode nothing could predict earns exactly what a correct prediction earns, and abstention credit is never printed without the false-alarm rate beside it, because a one-sided abstention number is unfalsifiable. Five tasks, ten null controls, sixteen refusal rules, no cross-task aggregate. The pixel column exists under a header reading NOT RANKED for one reason: so a control that replays ground-truth frames can top it outright and still lose the ranking. Two things a benchmark of this shape gets wrong by default, both of which this one got wrong first and had to be measured out: Unknowable episodes must be constructed, never asserted. The first version asserted them, and at S1 horizon 1 all 80 of 80 episodes labelled unknowable had a terminal state that was byte-identical for every value of the hidden latent, because that latent only acted from step 1. Fifty to seventy-nine percent of S5 and nine to thirty-six percent of S3 were the same. The abstention metrics were measuring the labeller. Every unknowable episode is now advanced under the full hidden grid at the horizon it is scored at, and must separate by more than two tolerances pairwise and one from nominal, or the pack build raises and emits nothing. Oracle success is evidence of knowability; oracle failure is not evidence of unknowability, and only the conclusive direction is read. A must-fail control that has never been observed to fail is not a control. N8 was differenced against a copy of itself and printed exactly +0.0000 on every task, forever, on any pack however leaky. It now demonstrates its probe works on the half it memorised before claiming anything about the half it did not, and it fails on a deliberately leaky pack, which is the only evidence that it can. The first scorecard is bad and that is the point. The built-in analytic floor prints NO INFORMATION on fourteen of sixteen rows, is rankable on two, and scores RefusalRecall 0.0000 almost everywhere: it never abstains on an episode nothing could predict, and it has no channel that could. Publishing that under the harness's own rules is the credibility event. A good-looking first row would have been the warning sign. The honest ceiling is in the README rather than implied. A world model predicts the visible half forward and is structurally blind to what the reagent actually is. Nothing here runs usefully CPU-only except the harness, nothing here belongs in a servo loop, and every row reads basis: simulated until a hardware run record exists.
The trigger was push on [main] plus pull_request, which reads as full coverage and is not. A branch whose pull request targets another branch can be reviewed and merged with this workflow never having run on it. Found by watching it happen: the first push of this branch ran house-rules and nothing else, and a silent job is harder to notice than a red one.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #2, so the diff here is only Phase 4 and the benchmark.
What this claims
Every world-model benchmark scores a rollout that was produced. This one scores the decision not to produce one: a correct refusal on an episode nothing could predict earns exactly what a correct prediction earns, no score is published for a model that ignores its action argument, and the same frozen threshold is what an arm path would gate on.
That gap is real as of Aug 2026, not assumed. A survey of roughly fifty world-model benchmarks (arXiv:2606.15032) names uncertainty calibration as one of seven required decision axes and finds it in one row. The closest existing work, RoboTrustBench (2606.01600), carries one refusal criterion of thirteen and, when a model actually refused to generate, excluded those samples as missing data rather than scoring them as correct abstentions.
What is in it
plr_lr/learn/-- Phase 4: record, model, plan, evaluate. CPU-only, no weights, no arm.plr_lr/learn/predict.py-- the contract a world model implements, carrying frames with hashes and capture times, plus the interventional gate that refuses rather than low-scores a model ignoring its actions.plr_lr/learn/screen.py-- admit or defer, with deferral terminal and no fallback action.plr_lr/learn/selective.py-- stdlib-only scoring core.scripts/score_standalone.pyis a single-file copy that runs on a barepython:3.9-slimwith zero installs, verified in CI.plr_lr/learn/bench/-- DeferLab: 5 tasks, 10 null controls, 16 refusal rules, no cross-task aggregate.Four measured defects fixed on the way in
DynamicsFit.improvedon pure noiseNonewith SPLIT-UNSOUNDReachTask5 trials at default jittervalid=TruevalidSamplingPlannerat n=256n_distinctKinestheticSessionon an injected trajectorysource=kinestheticpose_sourcerecords which it wasTwo traps this got wrong first and had to measure out
Unknowable labels must be constructed, not asserted. The first version asserted them, and at S1 horizon 1 all 80 of 80 episodes labelled unknowable were fully determined, because the hidden latent only acted from step 1. Also 50-79% of S5 and 9-36% of S3. The abstention metrics were measuring the labeller. Every unknowable episode is now advanced under the full hidden grid at its scored horizon and must separate pairwise by more than two tolerances, or the pack build raises and emits nothing.
A must-fail control never observed to fail is not a control. N8 was differenced against a copy of itself and printed exactly
+0.0000on every task, on any pack however leaky. It now proves its probe works on the half it memorised before claiming anything about the half it did not, and it fails on a deliberately leaky pack.The first scorecard is bad, deliberately
The built-in analytic floor prints
NO INFORMATIONon 14 of 16 rows, is rankable on 2, and scoresRefusalRecall 0.0000almost everywhere: it never abstains on an unknowable episode and has no channel that could. Publishing that under the harness's own rules is the point. A good-looking first row would have been the warning sign.Honest ceiling
A world model predicts the visible half forward and is structurally blind to what the reagent actually is. Nothing here belongs in a servo loop: V-JEPA 2-AC is 16 s per action on a 4090 and NVIDIA's own world-action-model number is 590-800 ms per chunk, against a 20 ms control period. Every row reads
basis: simulated; no arm has been run.Verification
1065 tests pass, ruff clean, house-rules clean over the full history, all four examples exit 0, and the standalone scorer produces a full scorecard on stock Python 3.9.6 with zero installs.