AI Potluck
Back to Gap Map Model components / Multimodal models

SmolVLM

Hugging Face
open weights / Overall score: 3.3

SmolVLM is Hugging Face's family of small vision-language models, from 256M to 2.2B parameters, built on the Idefics3 architecture for on-device use. The models caption, answer questions about and transcribe text from interleaved images, and SmolVLM2 adds video. Hugging Face publishes the training code in its smollm repository.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
partial(SmolVLM trained on The Cauldron and Docmatix, which Hugging Face publishes
code
open(pretraining and fine-tuning code in huggingface/smollm (vision/m4))
license
Apache-2.0(OSI

SmolVLM2 is Apache 2.0 and Hugging Face publishes its training code. The first SmolVLM trained on datasets Hugging Face itself releases, but SmolVLM2's video mixture is assembled from ten third-party sets and is listed rather than published.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across Hugging Face's SmolVLM and SmolVLM2 checkpoints. More than half of it is the 500M video model, which suits small on-device pipelines.

Capability

3 high confidence
3.0

SmolVLM2 watches video as well as reading images and documents, at sizes small enough for phones and laptops. Its scores are well below larger models, but it covers the same kinds of input as MiniCPM-V, so it sits level with it.

Verified 2026-09-26