Skip to content

Clarification on Stage III training data completeness #14

Description

@TossherO

Hi! Thanks for open-sourcing this great work.

I’m trying to reproduce Stage III training and want to confirm whether the currently released dataset is complete for that stage.

From my local analysis of the provided minecraft-vla-sft dataset, I currently see a single-image, single-action style dataset with about:

  • train: 3,776,475 samples.
  • valid: 1,000 samples.

However, in the paper’s data description I understood Stage III to involve broader trajectory data, including:

  • Human + agent gameplay trajectories (OpenAI contractor/VPT base data + additional VPT/JARVIS-1 rollout data, mentioned as extra millions of frames).
  • Synthetic GUI expert data (e.g., crafting/smelting), with millions of expert entries.
  • ~10B tokens used for Stage III pretraining, then final task-specific finetuning with held-out random trajectories.

Could you please clarify:

  • Is the currently released minecraft-vla-sft dataset only a subset of the full Stage III data?
  • Are the other trajectory / GUI expert datasets available publicly? If yes, where can they be downloaded? If they are not public, what is the recommended way to reproduce results with open data?
  • Are there expected schema/config differences when adding those additional datasets to the training pipeline? (The code appears to be designed for multi-frame image input, and the paper also includes models for multi-step action output.)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions