Hi! Thanks for open-sourcing this great work.
I’m trying to reproduce Stage III training and want to confirm whether the currently released dataset is complete for that stage.
From my local analysis of the provided minecraft-vla-sft dataset, I currently see a single-image, single-action style dataset with about:
- train: 3,776,475 samples.
- valid: 1,000 samples.
However, in the paper’s data description I understood Stage III to involve broader trajectory data, including:
- Human + agent gameplay trajectories (OpenAI contractor/VPT base data + additional VPT/JARVIS-1 rollout data, mentioned as extra millions of frames).
- Synthetic GUI expert data (e.g., crafting/smelting), with millions of expert entries.
- ~10B tokens used for Stage III pretraining, then final task-specific finetuning with held-out random trajectories.
Could you please clarify:
- Is the currently released minecraft-vla-sft dataset only a subset of the full Stage III data?
- Are the other trajectory / GUI expert datasets available publicly? If yes, where can they be downloaded? If they are not public, what is the recommended way to reproduce results with open data?
- Are there expected schema/config differences when adding those additional datasets to the training pipeline? (The code appears to be designed for multi-frame image input, and the paper also includes models for multi-step action output.)
Hi! Thanks for open-sourcing this great work.
I’m trying to reproduce Stage III training and want to confirm whether the currently released dataset is complete for that stage.
From my local analysis of the provided minecraft-vla-sft dataset, I currently see a single-image, single-action style dataset with about:
However, in the paper’s data description I understood Stage III to involve broader trajectory data, including:
Could you please clarify: