AI Potluck
Back to Gap Map Model components / Robotics & embodied AI

SmolVLA

Hugging Face
open source / Overall score: 2.0

SmolVLA is Hugging Face's compact vision-language-action model, about 450 million parameters, trained on community-collected datasets from affordable robot arms. It takes multi-view images, robot state and an optional instruction and outputs continuous action chunks by flow matching, trains on a single GPU, and runs on consumer GPUs or CPUs through the LeRobot library.

Openness

5 high confidence
5.0
weights
open
data
open(the paper releases the community training data)
code
open(training code in LeRobot)
license
Apache-2.0(weights and code)

Weights, training code and training data are all released, under Apache-2.0, so the model can be retrained from what is published.

Adoption

2 medium confidence
2.0

Hugging Face downloads of the base checkpoint. The benchmark fine-tunes published beside it are not declared here.

Capability

2 low confidence
2.0

SmolVLA is built for the low-cost arms its community data comes from, and its real-world results are on SO-100 and SO-101 arms. It learns from far less data than OpenVLA and does not claim control of robots outside that family, though its paper reports results close to models ten times its size.

  • https://arxiv.org/abs/2506.01844 recorded 2026-09-26

    Abstract: "Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10x larger"; evaluated "on a range of both simulated as well as real-world robotic benchmarks".

  • https://arxiv.org/html/2506.01844 recorded 2026-09-27

    Paper: "we selected a subset of 481 community datasets obtained from Hugging Face, filtered according to embodiment type, episode count, overall data quality, and frame coverage" (22.9K episodes, 10.6M frames); real-world datasets collected "using the SO-100 robot arm ... and 1 with SO-101 arm".

Verified 2026-09-26