Describe the bug
We fine-tuned WALL-X on the clean RoboTwin 2.0 demonstrations and evaluated the resulting checkpoint using the
official RoboTwin task definitions, seed manifests, and CuRobo planner.
The evaluation completes without runtime errors, but the complete 50-task result is substantially lower than expected:
demo_clean: 14.54% mean success rate
demo_randomized: 13.52% mean success rate
- 50 tasks
- 100 episodes per task/configuration
- 10,000 episodes in total
- All 100 task/configuration jobs completed
We would appreciate confirmation of the official WALL-X RoboTwin training and inference recipe, especially the action/
state representation, normalization, camera preprocessing, action execution horizon, and checkpoint selection.
-
Use WALL-X commit:
c74b619b99bb0a4e564f5f5e58f04592766841f0
-
Use RoboTwin commit:
c3ddfa8b97d5519efa828b075999bd0006778e5e
-
Fine-tune on RoboTwin 2.0 clean demonstrations:
- 50 tasks
- 50 demonstrations per task
- 2,500 episodes total
- LeRobot v3 format
- Evaluated checkpoint: step 60,000
- 8 GPUs
- Batch size per GPU: 4
- Gradient accumulation steps: 4
- Global batch size: 128
- Learning rate:
5e-5
- Minimum learning rate:
1e-6
- Warmup: 1,000 steps
- Weight decay:
1e-8
- Training seed: 10222
- Action horizon: 32
-
Use these dataset mappings:
action: action
state: observation.state
cameras:
observation.images.cam_high: face_view
observation.images.cam_left_wrist: left_wrist_view
observation.images.cam_right_wrist: right_wrist_view
-
Use a 14-dimensional physical state/action representation with 12 padded dimensions expected by the model
configuration.
-
Evaluate using:
- Official RoboTwin 2.0 50 tasks
- demo_clean and demo_randomized
- 100 episodes per task/configuration
- Official seed manifests
- Complete-manifest validation
- Official CuRobo planner
- Action horizon: 32
- Replan/action execution interval: 32
- Checkpoint-local normalization/configuration
-
Observed results:
Configuration Tasks Episodes/task Mean success
━━━━━━━━━━━━━━━━━ ━━━━━━━ ━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━
demo_clean 50 100 14.54%
───────────────── ─────── ─────────────── ──────────────
demo_randomized 50 100 13.52%
Stronger demo_clean task results include:
- shake_bottle_horizontally: 89%
- shake_bottle: 88%
- click_alarmclock: 68%
- click_bell: 65%
- adjust_bottle: 61%
- press_stapler: 53%
- grab_roller: 51%
Fourteen of the 50 tasks have zero success under demo_clean.
Expected behavior
We expected the official WALL-X fine-tuning and inference procedure to produce materially higher RoboTwin 2.0 success
rates.
Could you please clarify:
- Is there an official WALL-X RoboTwin 2.0 fine-tuning config or downstream checkpoint for numerical comparison?
- Is the 14-DoF state/action representation with 12 padded dimensions correct?
- Are the three camera mappings above correct?
- Should a complete 32-step action chunk be executed, or should the policy replan after fewer steps?
- Are checkpoint-local normalization statistics sufficient, or are separate RoboTwin-specific statistics required?
- Is step 60,000 at global batch size 128 in the expected training range for 2,500 clean episodes?
- Are RoboTwin-specific changes after commit c74b619... required for reproduction?
Environment
- OS: Linux x86_64
- Python: 3.10.12
- PyTorch: 2.7.1+cu128
- CUDA: 12.8
- GPU architecture: NVIDIA Hopper, compute capability 9.0
- WALL-X commit: c74b619b99bb0a4e564f5f5e58f04592766841f0
- RoboTwin commit: c3ddfa8b97d5519efa828b075999bd0006778e5e
Logs
There is no runtime exception. All evaluation jobs complete successfully, but the aggregate success rates are
unexpectedly low.
The complete sanitized training configuration, evaluation protocol, and per-task results are available and can be
provided if needed.
Describe the bug
We fine-tuned WALL-X on the clean RoboTwin 2.0 demonstrations and evaluated the resulting checkpoint using the
official RoboTwin task definitions, seed manifests, and CuRobo planner.
The evaluation completes without runtime errors, but the complete 50-task result is substantially lower than expected:
demo_clean: 14.54% mean success ratedemo_randomized: 13.52% mean success rateWe would appreciate confirmation of the official WALL-X RoboTwin training and inference recipe, especially the action/
state representation, normalization, camera preprocessing, action execution horizon, and checkpoint selection.
Use WALL-X commit:
c74b619b99bb0a4e564f5f5e58f04592766841f0Use RoboTwin commit:
c3ddfa8b97d5519efa828b075999bd0006778e5eFine-tune on RoboTwin 2.0 clean demonstrations:
5e-51e-61e-8Use these dataset mappings:
Use a 14-dimensional physical state/action representation with 12 padded dimensions expected by the model
configuration.
Evaluate using:
Observed results:
Configuration Tasks Episodes/task Mean success
━━━━━━━━━━━━━━━━━ ━━━━━━━ ━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━
demo_clean 50 100 14.54%
───────────────── ─────── ─────────────── ──────────────
demo_randomized 50 100 13.52%
Stronger demo_clean task results include:
Fourteen of the 50 tasks have zero success under demo_clean.
Expected behavior
We expected the official WALL-X fine-tuning and inference procedure to produce materially higher RoboTwin 2.0 success
rates.
Could you please clarify:
Environment
Logs
There is no runtime exception. All evaluation jobs complete successfully, but the aggregate success rates are
unexpectedly low.
The complete sanitized training configuration, evaluation protocol, and per-task results are available and can be
provided if needed.