Skip to content

[Bug] Low RoboTwin 2.0 success rate after WALL-X clean-data fine-tuning #116

Description

@zhanghaoyu913-cmyk

Describe the bug

We fine-tuned WALL-X on the clean RoboTwin 2.0 demonstrations and evaluated the resulting checkpoint using the
official RoboTwin task definitions, seed manifests, and CuRobo planner.

The evaluation completes without runtime errors, but the complete 50-task result is substantially lower than expected:

  • demo_clean: 14.54% mean success rate
  • demo_randomized: 13.52% mean success rate
  • 50 tasks
  • 100 episodes per task/configuration
  • 10,000 episodes in total
  • All 100 task/configuration jobs completed

We would appreciate confirmation of the official WALL-X RoboTwin training and inference recipe, especially the action/
state representation, normalization, camera preprocessing, action execution horizon, and checkpoint selection.

  1. Use WALL-X commit:

    c74b619b99bb0a4e564f5f5e58f04592766841f0

  2. Use RoboTwin commit:

    c3ddfa8b97d5519efa828b075999bd0006778e5e

  3. Fine-tune on RoboTwin 2.0 clean demonstrations:

    • 50 tasks
    • 50 demonstrations per task
    • 2,500 episodes total
    • LeRobot v3 format
    • Evaluated checkpoint: step 60,000
    • 8 GPUs
    • Batch size per GPU: 4
    • Gradient accumulation steps: 4
    • Global batch size: 128
    • Learning rate: 5e-5
    • Minimum learning rate: 1e-6
    • Warmup: 1,000 steps
    • Weight decay: 1e-8
    • Training seed: 10222
    • Action horizon: 32
  4. Use these dataset mappings:

    action: action
    state: observation.state
    cameras:
      observation.images.cam_high: face_view
      observation.images.cam_left_wrist: left_wrist_view
      observation.images.cam_right_wrist: right_wrist_view
    
  5. Use a 14-dimensional physical state/action representation with 12 padded dimensions expected by the model
    configuration.

  6. Evaluate using:

    • Official RoboTwin 2.0 50 tasks
    • demo_clean and demo_randomized
    • 100 episodes per task/configuration
    • Official seed manifests
    • Complete-manifest validation
    • Official CuRobo planner
    • Action horizon: 32
    • Replan/action execution interval: 32
    • Checkpoint-local normalization/configuration
  7. Observed results:

    Configuration Tasks Episodes/task Mean success
    ━━━━━━━━━━━━━━━━━ ━━━━━━━ ━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━
    demo_clean 50 100 14.54%
    ───────────────── ─────── ─────────────── ──────────────
    demo_randomized 50 100 13.52%

    Stronger demo_clean task results include:

    • shake_bottle_horizontally: 89%
    • shake_bottle: 88%
    • click_alarmclock: 68%
    • click_bell: 65%
    • adjust_bottle: 61%
    • press_stapler: 53%
    • grab_roller: 51%

    Fourteen of the 50 tasks have zero success under demo_clean.

Expected behavior

We expected the official WALL-X fine-tuning and inference procedure to produce materially higher RoboTwin 2.0 success
rates.

Could you please clarify:

  1. Is there an official WALL-X RoboTwin 2.0 fine-tuning config or downstream checkpoint for numerical comparison?
  2. Is the 14-DoF state/action representation with 12 padded dimensions correct?
  3. Are the three camera mappings above correct?
  4. Should a complete 32-step action chunk be executed, or should the policy replan after fewer steps?
  5. Are checkpoint-local normalization statistics sufficient, or are separate RoboTwin-specific statistics required?
  6. Is step 60,000 at global batch size 128 in the expected training range for 2,500 clean episodes?
  7. Are RoboTwin-specific changes after commit c74b619... required for reproduction?

Environment

  • OS: Linux x86_64
  • Python: 3.10.12
  • PyTorch: 2.7.1+cu128
  • CUDA: 12.8
  • GPU architecture: NVIDIA Hopper, compute capability 9.0
  • WALL-X commit: c74b619b99bb0a4e564f5f5e58f04592766841f0
  • RoboTwin commit: c3ddfa8b97d5519efa828b075999bd0006778e5e

Logs

There is no runtime exception. All evaluation jobs complete successfully, but the aggregate success rates are
unexpectedly low.

The complete sanitized training configuration, evaluation protocol, and per-task results are available and can be
provided if needed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions