AI Potluck
Back to Gap Map Model components / Robotics & embodied AI

OpenVLA

OpenVLA
open weights / Overall score: 3.0

OpenVLA is a seven-billion-parameter open vision-language-action model. It fine-tunes the Prismatic VLM (DINOv2 and SigLIP vision encoders with a Llama 2 language model) on 970k manipulation episodes from Open X-Embodiment, and controls robot arms seen in that mixture from a camera image and a language instruction.

GitHub marks the openvla repository as a fork of TRI-ML/prismatic-vlms, the codebase it grew from; it is OpenVLA's own repository. It has had no commits since March 2025.

Openness

3 medium confidence
3.0
weights
open
data
open(the Open X-Embodiment mixture)
code
open(pretraining, training and fine-tuning scripts)
license
MIT(checkpoints and code, per the card)+Llama-2-Community-License(the Llama 2 language model inside the weights)

The data, the training code and the checkpoints are all public, and the project releases them under MIT. The language model inside the weights is Meta's Llama 2, though, and Llama 2 comes under Meta's community license, which caps use by very large services.

Adoption

3 medium confidence
3.0

Hugging Face downloads of openvla-7b. Its LIBERO fine-tunes are downloaded separately and are not declared here.

Capability

3 medium confidence
3.0

OpenVLA learned from many robots in Open X-Embodiment and controls the single arms it saw there, but its own card says it does not carry over to a new robot without fine-tuning, and it offers nothing for bimanual or humanoid control.

  • https://huggingface.co/openvla/openvla-7b/raw/main/README.md recorded 2026-09-26

    Card: "can be used zero-shot to control robots for specific combinations of embodiments and domains seen in the Open-X pretraining mixture (e.g., for BridgeV2 environments with a Widow-X robot)"; "do not zero-shot generalize to new (unseen) robot embodiments".

Verified 2026-09-26