Hey,
Congrats on the great work! I had two questions:
-
Can the 467k caption/QA video-paired data that was used for training your model be released? It would help save time in running the data curation pipeline.
-
For evaluation results reported in Table 1, at what fps/number of frames was each video sampled at?
Also, a suggestion: in your HuggingFace dataset space, renaming "hands_eval" to "daily_eval" might be a good idea, as the paper reports by the split name "Daily".
Best,
Pulkit
Hey,
Congrats on the great work! I had two questions:
Can the 467k caption/QA video-paired data that was used for training your model be released? It would help save time in running the data curation pipeline.
For evaluation results reported in Table 1, at what fps/number of frames was each video sampled at?
Also, a suggestion: in your HuggingFace dataset space, renaming "hands_eval" to "daily_eval" might be a good idea, as the paper reports by the split name "Daily".
Best,
Pulkit