SmolVLM
Hugging FaceSmolVLM is Hugging Face's family of small vision-language models, from 256M to 2.2B parameters, built on the Idefics3 architecture for on-device use. The models caption, answer questions about and transcribe text from interleaved images, and SmolVLM2 adds video. Hugging Face publishes the training code in its smollm repository.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- partial(SmolVLM trained on The Cauldron and Docmatix, which Hugging Face publishes
- code
- open(pretraining and fine-tuning code in huggingface/smollm (vision/m4))
- license
- Apache-2.0(OSI
SmolVLM2 is Apache 2.0 and Hugging Face publishes its training code. The first SmolVLM trained on datasets Hugging Face itself releases, but SmolVLM2's video mixture is assembled from ten third-party sets and is listed rather than published.
- https://huggingface.co/api/models/HuggingFaceTB/SmolVLM2-2.2B-Instruct?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for SmolVLM2-2.2B-Instruct: gated false, two safetensors shards, card license apache-2.0
- https://huggingface.co/blog/smolvlm recorded 2026-09-26
"we used the datasets previously used for Idefics3: the Cauldron and Docmatix, which are also fully open-source"
- https://huggingface.co/HuggingFaceTB/SmolVLM2-2.2B-Instruct/raw/main/README.md recorded 2026-09-26
"SmolVLM2 used 3.3M samples for training originally from ten different datasets" with per-modality shares; "We release the SmolVLM2 checkpoints under the Apache 2.0 license"
- https://raw.githubusercontent.com/huggingface/smollm/HEAD/vision/m4/training/main.py recorded 2026-09-26
The m4 training entry point
- https://raw.githubusercontent.com/huggingface/smollm/HEAD/vision/README.md recorded 2026-09-26
"everything related to the training of our Vision Language Models series: SmolVLM. This includes pretraining and finetuning code"
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across Hugging Face's SmolVLM and SmolVLM2 checkpoints. More than half of it is the 500M video model, which suits small on-device pipelines.
- https://huggingface.co/api/models?author=HuggingFaceTB&search=SmolVLM&limit=100 recorded 2026-09-26
Twelve SmolVLM checkpoints (the app-config repository excluded) with 2,037,669 downloads in the trailing 30 days; SmolVLM2-500M-Video-Instruct 1,195,310
Capability
3 high confidenceSmolVLM2 watches video as well as reading images and documents, at sizes small enough for phones and laptops. Its scores are well below larger models, but it covers the same kinds of input as MiniCPM-V, so it sits level with it.
- https://huggingface.co/blog/smolvlm2 recorded 2026-09-26
"SmolVLM can now watch"; 256M, 500M and 2.2B models, "the smallest video language models ever released"
- https://huggingface.co/HuggingFaceTB/SmolVLM2-2.2B-Instruct/raw/main/README.md recorded 2026-09-26
"Multi-modal model (image/multi-image/video/text)"; "does not support image or video generation"; Video-MME 52.1, MMMU 42, OCRBench 72.9
Verified 2026-09-26