Skip to content

investigate and likely switch to openvla-oft #4

Description

@mehhl

there is a new version of OpenVLA, called OpenVLA-OFT link which claims SOTA (as in, better than pi0) performance on variousb enchmarks thanks to enhanced finetuning. The enhancements go far enough to significantly change the model architecture (up to 3 video streams as input, proprio as input, continuous actions rather than tokenization, ...) and the authors are 3 of the original OpenVLA paper authors, one of whom is also a co-founder of Physical Intelligence. It would seem prudent to follow this program:

  1. See if OpenVLA-OFT gives some meaningful improvement on our own tasks (nomagic-simple-box, etc.)
  2. If the answer to 1. is anything like "yes", switch any/all OpenVLA experiments to OpenVLA-OFT asap.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions