Skip to content

Prepare EgoScale Stage 1 DLC full training - #1

Draft
Apexcmgd wants to merge 65 commits into
mainfrom
dev/shengjunhe
Draft

Apexcmgd wants to merge 65 commits into
mainfrom
dev/shengjunhe

Conversation

@Apexcmgd

@Apexcmgd Apexcmgd commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • publish the finalized self-collected Stage 2 data-loading contract
  • add the six-dataset egoscale_stage1_ego_cartesian_clean_rl2 normalization and unified-action metadata
  • add a non-destructive Stage 1 DLC production wrapper with explicit 100k-step training defaults
  • add preflight-only support and update the staged-training runbook

Why

A fresh DLC checkout previously lacked the Stage 1 RL2 norm assets. The generic stage launcher also defaults to smoke settings, including 20 training steps and save/eval every 10 steps, which is unsafe for a full run unless every production value is overridden manually.

Validation

  • bash -n scripts/run_egoscale_stage.sh
  • bash -n scripts/run_egoscale_stage1_dlc_full.sh
  • 85 passed for tests/cotrain/test_action_space.py and tests/cotrain/test_rlds_dataset.py
  • clean DSW worktree at 26a13587a3a055183f6ce71883492513e2f68e2a
  • Stage 1 preflight passed for all six EgoVerse/EgoVerse-RL2 builders, all norm/mapping metadata, and /data/junhe/models/pi05_base/params
  • all 13 imported asset files match the verified DSW copies by SHA256; norm values are finite

Operational notes

No DLC GPU job was started by this PR. The runbook requires a separate 2-node x 8-GPU, 100-step checkpoint-and-resume validation before the 100k run. W&B credentials remain external and must be injected through DLC secrets.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants