Thank you for releasing the code and the Spatial-TTT-Data-97k dataset. I am trying to reproduce the training using the provided spatial_ttt_train.sh script.
I noticed that the README mentions training with Spatial-TTT-Data-97k, lact_chunk_size=2648, window_size=2648, video_max_frames=128, and 8 GPUs by default. However, I could not find detailed information about the expected training time and compute requirements.
Could you please share some details about the training resources used for the released/reproduction model?
Specifically, I would like to ask:
- What GPU type and number of GPUs were used for training?
- Approximately how long does training on Spatial-TTT-Data-97k take, in wall-clock time or GPU hours?
Thank you for releasing the code and the Spatial-TTT-Data-97k dataset. I am trying to reproduce the training using the provided spatial_ttt_train.sh script.
I noticed that the README mentions training with Spatial-TTT-Data-97k, lact_chunk_size=2648, window_size=2648, video_max_frames=128, and 8 GPUs by default. However, I could not find detailed information about the expected training time and compute requirements.
Could you please share some details about the training resources used for the released/reproduction model?
Specifically, I would like to ask: