diff --git a/CHANGELOG.md b/CHANGELOG.md index 6f5340a..9648d48 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,18 +1,15 @@ # Changelog -## 0.11.0 — 2026-09-14 - -- Add Scene Lab: CC0 captured geometry, 25,000 mesh-derived surface Gaussians, separate collision heightfield, two actual L40S OptiX renders, bounded edit recipes and independent native Blender readback. -- Add the official LIBERO-Plus adapter and source-linked SmolVLA replay: paired baseline, camera viewpoint ID 609 and lighting ID 2124. Outcomes are success at 77 actions, step limit at 220, and success at 87. This is a diagnostic subset, not a benchmark score. -- Add an OpenEnv 0.4.2 verifiable creation environment. Numeric actions build and independently reopen a native scene; completion claims do not earn reward. -- Add a pinned Cosmos Policy recorder/access preflight. Actual Cosmos inference remains unavailable until the account receives access to its gated NVIDIA dependency. No prediction results are claimed. -- Link all 27 cross-project skill trials from the homepage and Hugging Face Space; improve mobile layouts, video seeking and lazy Gaussian loading. +## Unreleased +## 0.12.0 — 2026-09-14 -## Unreleased +- Inspect failure classifications in Stress Lab and jump to each recorded final motion window without dropping experiment totals. Show measured travel in millimetres and keep causal limits explicit. +- Accept original released stress packs that predate `reliability.json`; new packs still seal and independently recompute it. +- Reject identical-input claims when a repeat lacks consumed frames or both records lack camera hashes. -- Say what a failure was, not only that it happened: split step-limit failures by whether the arm was still moving at the cut-off. All 14 recorded failures were still in motion, so the 160-action budget is binding on the reported success rate rather than the policy having given up. -- Measure whether a repeat of the identical plan agrees, on the same L40S, at every level: 30/30 outcomes, action counts, bitwise physics states and actions, and 360/360 renders on the frames a policy call consumed. Report renders on consumed frames apart from recorded-only frames, because only the former can change an outcome. One recorded-only frame of 3,195 differed and did not propagate. +- Say what a failure was, not only that it happened: split step-limit failures by whether the arm was still moving at the cut-off. All 14 recorded failures were still in motion, under the stated motion threshold; continued movement does not establish progress or success with a larger action budget. +- Measure whether a repeat of the identical plan agrees, on the same L40S, at every level: 30/30 outcomes, action counts, numerically equal recorded robot states and actions, and 360/360 renders on the frames a policy call consumed. Report renders on consumed frames apart from recorded-only frames, because only the former can change an outcome. One recorded-only frame of 3,195 differed and did not propagate. - Seal the failure taxonomy inside the experiment pack and recompute it on verification, rather than only checking its hash. Keep the reproducibility aggregate beside the pack, since it is a property of two collections and belongs to neither. - Verify draft assets against local checksums and GitHub SHA-256 digests, retry @@ -20,6 +17,16 @@ Preserve existing public releases and drafts; document recovery using the original tested CI artifacts. Record the completed 0.10.0 publication. +## 0.11.0 — 2026-09-14 + +- Add Scene Lab: CC0 captured geometry, 25,000 mesh-derived surface Gaussians, separate collision heightfield, two actual L40S OptiX renders, bounded edit recipes and independent native Blender readback. +- Add the official LIBERO-Plus adapter and source-linked SmolVLA replay: paired baseline, camera viewpoint ID 609 and lighting ID 2124. Outcomes are success at 77 actions, step limit at 220, and success at 87. This is a diagnostic subset, not a benchmark score. +- Add an OpenEnv 0.4.2 verifiable creation environment. Numeric actions build and independently reopen a native scene; completion claims do not earn reward. +- Add a pinned Cosmos Policy recorder/access preflight. Actual Cosmos inference remains unavailable until the account receives access to its gated NVIDIA dependency. No prediction results are claimed. +- Link all 27 cross-project skill trials from the homepage and Hugging Face Space; improve mobile layouts, video seeking and lazy Gaussian loading. + + + ## 0.10.0 — One launch, six recorded futures - Add Solver Lab: six independent L40S/CUDA ballistic flights from Genesis 1.4.1 diff --git a/CITATION.cff b/CITATION.cff index 9b8fba2..e12a617 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -7,7 +7,7 @@ authors: repository-code: "https://github.com/noteflowai/robot-reel" url: "https://noteflowai.github.io/robot-reel/" license: Apache-2.0 -version: 0.11.0 +version: 0.12.0 date-released: 2026-09-14 abstract: "Robot Reel is a physical-AI replay lab that records MuJoCo and Newton physics and SmolVLA/LeRobot rollouts, verifies the recorded traces, and directs them into inspectable browser replays and editable Blender/OpenUSD films through a CLI and an MCP agent director." keywords: diff --git a/README.md b/README.md index 7525d25..6668cc4 100644 --- a/README.md +++ b/README.md @@ -190,17 +190,21 @@ report for independent verification with the 0.8.0+ installed CLI. [Compare paired outcomes](https://noteflowai.github.io/robot-reel/stress/#outcomes) · [Report method and CLI](docs/stress.md#compare-paired-outcomes). +**[Inspect every failure](https://noteflowai.github.io/robot-reel/stress/#failures).** Filter the recorded classifications and jump to each final motion window, with measured end-effector travel. + +[![Recorded failure review: classified episodes, complete denominators and a jump to the final motion window.](docs/failure-review.png)](https://noteflowai.github.io/robot-reel/stress/#failures) + **What the failures were.** A success rate does not say. All 14 failures here end at the step limit, and that covers a policy that froze and one still reaching when the budget expired. Measured: **every one was still in motion at the cut-off**, 44.8 mm to 138.6 mm of end-effector travel over the final tenth of its episode, -against a 1 mm stall threshold. So 160 actions is binding on the reported success -rate, not a policy that gave up. +against a 1 mm stall threshold. The recorded budget expired while the arm was moving. This does not show +that extra actions would complete the task or that the motion made progress. **Whether a repeat agrees.** The paired groups blame a condition for an outcome flip, which only holds if the same seed and condition answer the same twice. The whole plan was re-run on the same L40S: **30 / 30 identical outcomes, action -counts, physics states and actions**, and **360 / 360 identical renders on the +counts, recorded robot states and actions**, and **360 / 360 identical renders on the frames a policy call consumed**. One of 3,195 recorded-only frames differed, on a frame no call consumed. That difference is reported rather than rounded away: it shows the renderer is not bitwise deterministic even on identical hardware, and it diff --git a/README.zh-CN.md b/README.zh-CN.md index 70cac23..3a1420d 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -174,14 +174,18 @@ Markdown,重新导入可恢复原始样本位置,也可用命令行与完整 记录明确标注末帧停留,保留完整实验统计,个人备注与录制事实分开。 [复盘流程与核验命令](docs/stress.md#share-a-moment-for-review)。 +**[逐次检查失败](https://noteflowai.github.io/robot-reel/stress/#failures)。** 按记录类别筛选,直接跳到每个回合的末段运动窗口,并查看以毫米表示的末端移动距离。 + +[![真实失败分类与末段运动复盘界面](docs/failure-review.png)](https://noteflowai.github.io/robot-reel/stress/#failures) + **失败究竟是什么。** 成功率不回答这个问题。这里 14 次失败全部止于步数上限,而这个标签 同时涵盖「策略停住了」和「预算耗尽时仍在伸手」两种情形。实测结果是:**每一次失败在截断时 都仍在运动**,末端在最后十分之一个回合内移动 44.8 mm 到 138.6 mm,而停滞阈值为 1 mm。 -所以 160 步的预算才是所报成功率的约束,而不是策略放弃了。 +记录预算耗尽时机械臂仍在运动,但这不能证明增加步数会成功,也不能证明该运动代表任务进展。 **重跑是否给出同样的答案。** 配对分组把结果翻转归因于条件,这只在同一种子与条件两次给出 -同样答案时才成立。整个计划在同一块 L40S 上重跑了一遍:**结果、动作步数、逐位物理状态与 -动作均为 30 / 30 一致**,**策略调用实际消费的帧渲染 360 / 360 一致**。3,195 个仅记录帧中 +同样答案时才成立。整个计划在同一块 L40S 上重跑了一遍:**结果、动作步数、记录中的机器人状态与 +动作数值均为 30 / 30 一致**,**策略调用实际消费的帧渲染 360 / 360 一致**。3,195 个仅记录帧中 有 1 帧不同,且该帧未被任何推理调用消费。这个差异被如实报告而非抹平:它说明即使在同一硬件上 渲染也不是逐位确定的,而它没有传播仅仅是因为落点位置。 [分类、可复现性及其局限](docs/stress.md#what-the-failures-were-and-whether-a-repeat-agrees)。 diff --git a/docs/failure-review.png b/docs/failure-review.png new file mode 100644 index 0000000..5a0f89a Binary files /dev/null and b/docs/failure-review.png differ diff --git a/docs/stress.md b/docs/stress.md index d99c45c..dcd251a 100644 --- a/docs/stress.md +++ b/docs/stress.md @@ -51,8 +51,8 @@ UI filter never removes pairs from the report. “Not completed” combines A success rate does not say what a failure was. All fourteen failures here end at the step limit, and that label covers two different things: a policy that stopped, -and a policy still reaching when the budget expired. They call for opposite -responses. +and an arm still moving when the budget expired. This measurement describes +the recorded motion; it does not identify which intervention would fix the task. ```bash python3 -m robot_reel.cli stress docs/stress --reliability @@ -61,9 +61,9 @@ python3 -m robot_reel.cli stress docs/stress --reliability Measured over the published collection, **every one of the fourteen failures was still in motion at the cut-off**; none had stalled. End-effector travel over the final tenth of each episode ranges from 44.8 mm to 138.6 mm, against a stall -threshold of 1 mm chosen to sit far below anything a moving arm produces. So the -step limit of 160 actions is binding on the reported success rate rather than the -policy having given up, and raising it would be expected to change that rate. +threshold of 1 mm. The 160-action budget expired while motion continued. +That does not establish progress toward the goal or that a larger budget would +change the success rate; either conclusion would require a new controlled run. Failures also wander further to end up no further along: median path-to- displacement ratio 3.17 against 1.92 for successes. The ranges overlap, so that @@ -87,16 +87,18 @@ python3 -m robot_reel.cli stress docs/stress --repeat artifacts/repro-gpu-30 | --- | --- | | Same outcome | 30 / 30 trials | | Same action count | 30 / 30 trials | -| Bitwise identical physics states | 30 / 30 trials | -| Bitwise identical actions | 30 / 30 trials | +| Numerically equal recorded robot states | 30 / 30 trials | +| Numerically equal recorded actions | 30 / 30 trials | | Identical renders on frames a policy call consumed | 360 / 360 frames | | Identical renders on frames recorded only | 3194 / 3195 frames | The two render levels are separated because only the first can change an outcome. The policy is called once every ten steps, so most recorded frames never reach it. The single differing frame is the wrist view at frame 3 of -`seed-05-camera`, which no policy call consumed; the physics states and actions -of that trial are bitwise identical throughout. +`seed-05-camera`, which no policy call consumed; the recorded robot states and actions +of that trial compare numerically equal throughout. This compares the saved +`state` and `action` arrays, not every internal MuJoCo variable or floating-point +bit pattern (for example, numeric equality treats signed zeros as equal). That difference is worth stating rather than rounding away. It shows the renderer is not bitwise deterministic across runs even on identical hardware, and it did @@ -383,3 +385,9 @@ the ordinary imageio-ffmpeg/Pillow dependencies. Neither check reruns a policy. See the included media notice for third-party attribution. Model weights and upstream simulation meshes are not redistributed in the public experiment. + +## Inspect final motion and retain legacy packs + +Stress Lab now groups the recorded outcomes and lets a reviewer jump to the final tenth of samples of any selected episode. It displays the exact sample window and end-effector path length, with a stated 1 mm threshold. Movement alone cannot establish progress or the effect of a larger action budget. The full matrix, denominators and attempt ledger remain available. + +The verifier accepts the original exact inventory of releases through v0.11.0, which predated `reliability.json`. New exports include the derived taxonomy and independently recompute it. A missing consumed frame, or absent camera hashes in both runs, cannot count as identical policy input. diff --git a/docs/stress/METHODS.md b/docs/stress/METHODS.md index d99c45c..dcd251a 100644 --- a/docs/stress/METHODS.md +++ b/docs/stress/METHODS.md @@ -51,8 +51,8 @@ UI filter never removes pairs from the report. “Not completed” combines A success rate does not say what a failure was. All fourteen failures here end at the step limit, and that label covers two different things: a policy that stopped, -and a policy still reaching when the budget expired. They call for opposite -responses. +and an arm still moving when the budget expired. This measurement describes +the recorded motion; it does not identify which intervention would fix the task. ```bash python3 -m robot_reel.cli stress docs/stress --reliability @@ -61,9 +61,9 @@ python3 -m robot_reel.cli stress docs/stress --reliability Measured over the published collection, **every one of the fourteen failures was still in motion at the cut-off**; none had stalled. End-effector travel over the final tenth of each episode ranges from 44.8 mm to 138.6 mm, against a stall -threshold of 1 mm chosen to sit far below anything a moving arm produces. So the -step limit of 160 actions is binding on the reported success rate rather than the -policy having given up, and raising it would be expected to change that rate. +threshold of 1 mm. The 160-action budget expired while motion continued. +That does not establish progress toward the goal or that a larger budget would +change the success rate; either conclusion would require a new controlled run. Failures also wander further to end up no further along: median path-to- displacement ratio 3.17 against 1.92 for successes. The ranges overlap, so that @@ -87,16 +87,18 @@ python3 -m robot_reel.cli stress docs/stress --repeat artifacts/repro-gpu-30 | --- | --- | | Same outcome | 30 / 30 trials | | Same action count | 30 / 30 trials | -| Bitwise identical physics states | 30 / 30 trials | -| Bitwise identical actions | 30 / 30 trials | +| Numerically equal recorded robot states | 30 / 30 trials | +| Numerically equal recorded actions | 30 / 30 trials | | Identical renders on frames a policy call consumed | 360 / 360 frames | | Identical renders on frames recorded only | 3194 / 3195 frames | The two render levels are separated because only the first can change an outcome. The policy is called once every ten steps, so most recorded frames never reach it. The single differing frame is the wrist view at frame 3 of -`seed-05-camera`, which no policy call consumed; the physics states and actions -of that trial are bitwise identical throughout. +`seed-05-camera`, which no policy call consumed; the recorded robot states and actions +of that trial compare numerically equal throughout. This compares the saved +`state` and `action` arrays, not every internal MuJoCo variable or floating-point +bit pattern (for example, numeric equality treats signed zeros as equal). That difference is worth stating rather than rounding away. It shows the renderer is not bitwise deterministic across runs even on identical hardware, and it did @@ -383,3 +385,9 @@ the ordinary imageio-ffmpeg/Pillow dependencies. Neither check reruns a policy. See the included media notice for third-party attribution. Model weights and upstream simulation meshes are not redistributed in the public experiment. + +## Inspect final motion and retain legacy packs + +Stress Lab now groups the recorded outcomes and lets a reviewer jump to the final tenth of samples of any selected episode. It displays the exact sample window and end-effector path length, with a stated 1 mm threshold. Movement alone cannot establish progress or the effect of a larger action budget. The full matrix, denominators and attempt ledger remain available. + +The verifier accepts the original exact inventory of releases through v0.11.0, which predated `reliability.json`. New exports include the derived taxonomy and independently recompute it. A missing consumed frame, or absent camera hashes in both runs, cannot count as identical policy input. diff --git a/docs/stress/index.html b/docs/stress/index.html index bfd3a2a..8412879 100644 --- a/docs/stress/index.html +++ b/docs/stress/index.html @@ -38,6 +38,14 @@

Change the view.
Check the policy.

“Not completed” includes step limits and terminated runs; execution errors remain in the attempt ledger. Groups use the complete set of paired seeds. These are observed outcomes on one task, not a significance test or a general robustness score.

Check an exported report independently

With Robot Reel 0.8.0+ installed: robot-reel stress stress-lab --paired-report paired-outcomes.json. Replace stress-lab with your extracted experiment folder. The command checks the full experiment before comparing every group and count. From a source checkout, use python3 -m robot_reel.cli stress docs/stress --paired-report paired-outcomes.json.

+
+

What happened at the cut-off?

+

+
+

+

Selected episode

Complete classification JSON ↓
+

Motion is measured over the final tenth of recorded samples, with a 1 mm stall threshold. It does not establish task progress, the cause of failure, or success with a larger action budget. Filtering keeps the experiment totals and attempt ledger intact.

+

Take the evidence with you.

Every camera, applied action, inference timing, initial state and result. The offline folder embeds its telemetry and needs no service.

-
Complete offline lab ↓MCAP telemetry ↓Results CSV ↓Summary JSON ↓Locked plan ↓Attempt ledger ↓
+
Complete offline lab ↓MCAP telemetry ↓Results CSV ↓Summary JSON ↓Locked plan ↓Attempt ledger ↓

MCAP contains JSON observation and inference channels, with episode-relative nanosecond timestamps. Videos are separate MP4s. No model weights are bundled.

Inspect every recorded result and source file
TrialOutcomeActionsEpisodePolicy callsEvidence
Attempt history
Trial / attemptStatusStarted UTCResult or error
@@ -99,6 +107,34 @@

Turn a moment into a review.

let pair=[],master=null,distances=[],peakFrame=0; let outcomeFilter='all',outcomeCondition=condition==='reference'?'dim':condition; function element(tag,text,className){const el=document.createElement(tag);el.textContent=text;if(className)el.className=className;return el;} +const failureLabels={success:'Success',step_limit_in_motion:'Step limit · still moving',step_limit_stalled:'Step limit · below motion threshold',terminated:'Terminated'}; +const failureRuns=new Map(data.traces.map(t=>{ + const points=t.frames.map(f=>f.state.slice(0,3)),start=points.length-Math.max(2,Math.floor(points.length/10)); + let travel=0;for(let i=start+1;iv-points[i-1][j])); + const kind=t.result.outcome==='step_limit'?(travel<.001?'step_limit_stalled':'step_limit_in_motion'):t.result.outcome; + return [t.stress.trial_id,{kind,travel,start}]; +})); +for(const kind of ['all',...Object.keys(failureLabels)]){ + const count=kind==='all'?failureRuns.size:[...failureRuns.values()].filter(r=>r.kind===kind).length; + const option=element('option',`${kind==='all'?'All completed trials':failureLabels[kind]} · ${count}`);option.value=kind;$('failure-kind').append(option); +} +$('failure-kind').value=[...failureRuns.values()].some(r=>r.kind==='step_limit_in_motion')?'step_limit_in_motion':'all'; +$('failure-summary').textContent=`${data.summary.completed_trials} recorded trials / ${data.summary.planned_trials} planned. Select a trial to inspect its final motion window alongside the paired reference.`; +function renderFailures(){ + const kind=$('failure-kind').value,box=$('failure-trials');box.replaceChildren(); + for(const t of data.traces){ + const run=failureRuns.get(t.stress.trial_id);if(kind!=='all'&&run.kind!==kind)continue; + const button=element('button',`Seed ${t.seed} · ${label(t.stress.condition)}`);button.type='button'; + button.dataset.trial=t.stress.trial_id;button.setAttribute('aria-pressed',String(seed===t.seed&&condition===t.stress.condition)); + button.onclick=()=>{seed=t.seed;condition=t.stress.condition;frame=run.start;loadPair();document.querySelector('.workbench').scrollIntoView({block:'start'});$('play').focus({preventScroll:true});}; + box.append(button); + } + $('failure-count').textContent=`${box.childElementCount} of ${failureRuns.size} recorded trials in this view.${box.childElementCount?'':' No recorded trial matches this classification.'}`; + const selected=get(seed,condition),run=failureRuns.get(selected.stress.trial_id); + $('failure-detail').textContent=`${selected.stress.trial_id} · ${failureLabels[run.kind]||run.kind}. ${(run.travel*1000).toFixed(1)} mm of end-effector travel across the last ${selected.frames.length-run.start} samples (${run.start}–${selected.frames.length-1}).`; +} +$('failure-kind').onchange=renderFailures; +$('failure-tail').onclick=()=>{selectFrame(failureRuns.get(get(seed,condition).stress.trial_id).start,'Inspecting the recorded final motion window.');document.querySelector('.workbench').scrollIntoView({block:'start'});$('play').focus({preventScroll:true});}; const outcomeLabels={both_success:'Success → success',lost_success:'Success → not completed',gained_success:'Not completed → success',neither_success:'Neither completed'}; function pairedOutcomes(){ return { @@ -241,7 +277,7 @@

Turn a moment into a review.

const curve=distances.map((v,n)=>`${n?'L':'M'}${n/(distances.length-1)*720},${70-v/high*66}`).join(' '); $('distance-line').setAttribute('d',curve);$('distance-area').setAttribute('d',curve+' L720,72 L0,72 Z'); $('distance-caption').textContent=`Peak ${(distances[peakFrame]*100).toFixed(1)} cm at sample ${peakFrame} · ${distances.length} shared observations. Measured trajectory difference, not a task-error or accuracy score.`; - master=videos[end(pair[0])>=end(pair[1])?0:1];render();renderOutcomeLens(); + master=videos[end(pair[0])>=end(pair[1])?0:1];render();renderOutcomeLens();renderFailures(); } function selectFrame(n,message){pause(message);frame=n;render();} function tick(){ diff --git a/docs/stress/manifest.json b/docs/stress/manifest.json index 7abe1a4..b3b2f0d 100644 --- a/docs/stress/manifest.json +++ b/docs/stress/manifest.json @@ -2,11 +2,11 @@ "schema": "robot-reel-stress-1", "files": { "LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4", - "METHODS.md": "af88bc74853f90dd3aa877f9043960c0076ed78efb88a030fba626f1d7d9aa93", + "METHODS.md": "cda076a5fe7614e8dada470e7f7abbe1e133a399c39cac0156e2a7d3383f4c26", "NOTICE.txt": "ea626872970c7acaf4faafb6fa4a904d2f7b885e376c9cf6ba249b71c3e5ad3e", "attempts.json": "ae3fc2cc21f3eb0808836eacb39ed26fdeedbb7ae36141927f2efde68676f938", "experiment.json": "d3b2292a818db5298e2b42562170ba8c263bb44c25f6265ad65473cd39b77bf0", - "index.html": "33168032866701b9253f92bef7f409071cb3647f8f8d39e3f2f34297d3f220c7", + "index.html": "6aaa7addeb1f501f2cf463f3fcfedbff207fdc400a2ff59ed735d773d4c3a3fd", "media-checks.json": "5019ec2eaf03244ea91cb1357d611cc364b80da0e13ae43401b6e83fd2521db0", "reliability.json": "4b1a189366d5e954762b0d8f2f89b7b0c36671474c52c8d0cfeecbc65cbdbc41", "results.csv": "fa7799110bb38edfb74dbbc5f6393574615d1d56e2a273873a3d6208af6a4126", diff --git a/pyproject.toml b/pyproject.toml index 11df7b5..be973e5 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "robot-reel" -version = "0.11.0" +version = "0.12.0" description = "Inspect real VLA rollouts and direct recorded physics into editable Blender/OpenUSD films." readme = "README.md" requires-python = ">=3.12" diff --git a/robot_reel/reliability.py b/robot_reel/reliability.py index c4e4f54..4e57963 100644 --- a/robot_reel/reliability.py +++ b/robot_reel/reliability.py @@ -2,8 +2,8 @@ A success rate answers neither question. Fourteen of thirty recorded trials end at the step limit, and that label alone does not distinguish a policy that froze -from one that was still reaching when the budget ran out. The two call for -opposite responses: raise the step limit, or fix the policy. +from one that was still moving when the budget ran out. This describes the +recording; selecting an effective intervention requires a controlled experiment. Reproducibility is the other half. The paired design attributes an outcome flip to a condition, which only holds if the same seed and condition give the same @@ -20,10 +20,9 @@ SCHEMA = "robot-reel-reliability-1" -# A tenth of a millimetre of end-effector travel over the final tenth of an -# episode is below anything a moving arm produces, so a run under it was not -# merely slow. The threshold is stated rather than fitted: the smallest tail -# motion in the recorded collection is 24.7 mm, three orders of magnitude above. +# One millimetre of end-effector travel over the final tenth of recorded frames. +# This descriptive threshold does not distinguish slow progress from a stall, +# nor establish that a moving arm would succeed with a larger action budget. STALL_METRES = 1e-3 @@ -44,6 +43,8 @@ def classify_run(trace, *, stall_metres=STALL_METRES): early termination is a different event from exhausting the budget. """ points = _positions(trace) + if not math.isfinite(stall_metres) or stall_metres <= 0: + raise ValueError("stall_metres must be finite and positive") if len(points) < 2: raise ValueError("A trial needs at least two recorded frames") outcome = trace["result"]["outcome"] @@ -106,27 +107,38 @@ def compare_run(reference, repeat, action_steps): """Compare one trial against its repeat, from outcome down to pixels. Reported as separate levels because they are separate claims. Agreeing - outcomes are what the paired comparison rests on; identical physics is a - stronger statement; identical renders on consumed frames is stronger again. + outcomes are what the paired comparison rests on; matching recorded robot + states and matching consumed images provide separate evidence. State/action + arrays are compared numerically, not as floating-point bit patterns or as + a complete serialization of the simulator. """ a, b = reference["frames"], repeat["frames"] action_steps = int(action_steps) if action_steps < 1: raise ValueError("action_steps must be positive") - consumed = _input_frames(a, action_steps) | _input_frames(b, action_steps) + inputs_a, inputs_b = _input_frames(a, action_steps), _input_frames(b, action_steps) + consumed = inputs_a | inputs_b shared = range(min(len(a), len(b))) render = {"input_frames": [0, 0], "recorded_frames": [0, 0]} for index in shared: key = "input_frames" if index in consumed else "recorded_frames" render[key][1] += 1 - render[key][0] += a[index].get("raw_camera_sha256") == b[index].get("raw_camera_sha256") + hashes_a, hashes_b = a[index].get("raw_camera_sha256"), b[index].get("raw_camera_sha256") + complete = all( + isinstance(hashes, dict) + and all(isinstance(hashes.get(camera), str) and hashes[camera] + for camera in ("main", "wrist")) + for hashes in (hashes_a, hashes_b) + ) + render[key][0] += complete and hashes_a == hashes_b return { "same_outcome": reference["result"]["outcome"] == repeat["result"]["outcome"], "same_action_count": reference["result"]["actions"] == repeat["result"]["actions"], "same_frame_count": len(a) == len(b), "identical_states": [frame["state"] for frame in a] == [frame["state"] for frame in b], "identical_actions": [frame.get("action") for frame in a] == [frame.get("action") for frame in b], - "identical_input_renders": render["input_frames"][0] == render["input_frames"][1], + "identical_input_renders": bool(consumed) and inputs_a == inputs_b + and render["input_frames"][0] == len(consumed), "input_renders": {"identical": render["input_frames"][0], "compared": render["input_frames"][1]}, "recorded_renders": {"identical": render["recorded_frames"][0], "compared": render["recorded_frames"][1]}, } diff --git a/robot_reel/resources/stress/METHODS.txt b/robot_reel/resources/stress/METHODS.txt index d99c45c..dcd251a 100644 --- a/robot_reel/resources/stress/METHODS.txt +++ b/robot_reel/resources/stress/METHODS.txt @@ -51,8 +51,8 @@ UI filter never removes pairs from the report. “Not completed” combines A success rate does not say what a failure was. All fourteen failures here end at the step limit, and that label covers two different things: a policy that stopped, -and a policy still reaching when the budget expired. They call for opposite -responses. +and an arm still moving when the budget expired. This measurement describes +the recorded motion; it does not identify which intervention would fix the task. ```bash python3 -m robot_reel.cli stress docs/stress --reliability @@ -61,9 +61,9 @@ python3 -m robot_reel.cli stress docs/stress --reliability Measured over the published collection, **every one of the fourteen failures was still in motion at the cut-off**; none had stalled. End-effector travel over the final tenth of each episode ranges from 44.8 mm to 138.6 mm, against a stall -threshold of 1 mm chosen to sit far below anything a moving arm produces. So the -step limit of 160 actions is binding on the reported success rate rather than the -policy having given up, and raising it would be expected to change that rate. +threshold of 1 mm. The 160-action budget expired while motion continued. +That does not establish progress toward the goal or that a larger budget would +change the success rate; either conclusion would require a new controlled run. Failures also wander further to end up no further along: median path-to- displacement ratio 3.17 against 1.92 for successes. The ranges overlap, so that @@ -87,16 +87,18 @@ python3 -m robot_reel.cli stress docs/stress --repeat artifacts/repro-gpu-30 | --- | --- | | Same outcome | 30 / 30 trials | | Same action count | 30 / 30 trials | -| Bitwise identical physics states | 30 / 30 trials | -| Bitwise identical actions | 30 / 30 trials | +| Numerically equal recorded robot states | 30 / 30 trials | +| Numerically equal recorded actions | 30 / 30 trials | | Identical renders on frames a policy call consumed | 360 / 360 frames | | Identical renders on frames recorded only | 3194 / 3195 frames | The two render levels are separated because only the first can change an outcome. The policy is called once every ten steps, so most recorded frames never reach it. The single differing frame is the wrist view at frame 3 of -`seed-05-camera`, which no policy call consumed; the physics states and actions -of that trial are bitwise identical throughout. +`seed-05-camera`, which no policy call consumed; the recorded robot states and actions +of that trial compare numerically equal throughout. This compares the saved +`state` and `action` arrays, not every internal MuJoCo variable or floating-point +bit pattern (for example, numeric equality treats signed zeros as equal). That difference is worth stating rather than rounding away. It shows the renderer is not bitwise deterministic across runs even on identical hardware, and it did @@ -383,3 +385,9 @@ the ordinary imageio-ffmpeg/Pillow dependencies. Neither check reruns a policy. See the included media notice for third-party attribution. Model weights and upstream simulation meshes are not redistributed in the public experiment. + +## Inspect final motion and retain legacy packs + +Stress Lab now groups the recorded outcomes and lets a reviewer jump to the final tenth of samples of any selected episode. It displays the exact sample window and end-effector path length, with a stated 1 mm threshold. Movement alone cannot establish progress or the effect of a larger action budget. The full matrix, denominators and attempt ledger remain available. + +The verifier accepts the original exact inventory of releases through v0.11.0, which predated `reliability.json`. New exports include the derived taxonomy and independently recompute it. A missing consumed frame, or absent camera hashes in both runs, cannot count as identical policy input. diff --git a/robot_reel/stress.html b/robot_reel/stress.html index 616ce3e..b2bc2ec 100644 --- a/robot_reel/stress.html +++ b/robot_reel/stress.html @@ -38,6 +38,14 @@

Change the view.
Check the policy.

“Not completed” includes step limits and terminated runs; execution errors remain in the attempt ledger. Groups use the complete set of paired seeds. These are observed outcomes on one task, not a significance test or a general robustness score.

Check an exported report independently

With Robot Reel 0.8.0+ installed: robot-reel stress stress-lab --paired-report paired-outcomes.json. Replace stress-lab with your extracted experiment folder. The command checks the full experiment before comparing every group and count. From a source checkout, use python3 -m robot_reel.cli stress docs/stress --paired-report paired-outcomes.json.

+
+

What happened at the cut-off?

+

+
+

+ +

Motion is measured over the final tenth of recorded samples, with a 1 mm stall threshold. It does not establish task progress, the cause of failure, or success with a larger action budget. Filtering keeps the experiment totals and attempt ledger intact.

+