Hi! When reading Appendix C, I ran into some confusion about how action chunks are implemented.
When I run inference using your open-sourced model, I notice that the model only outputs one set of actions per frame.
To produce an action chunk, do you:
feed the model’s previous output back into the model again through multiple rounds of prompting,
or
have the model generate multiple actions in one single forward pass?
If it’s the second case, then how does the model produce an action chunk of size 2 during inference, given that the action-chunk sizes used during training were 1 and 3?
Hi! When reading Appendix C, I ran into some confusion about how action chunks are implemented.
When I run inference using your open-sourced model, I notice that the model only outputs one set of actions per frame.
To produce an action chunk, do you:
feed the model’s previous output back into the model again through multiple rounds of prompting,
or
have the model generate multiple actions in one single forward pass?
If it’s the second case, then how does the model produce an action chunk of size 2 during inference, given that the action-chunk sizes used during training were 1 and 3?