Idefics
Hugging FaceIdefics is Hugging Face's line of open vision-language models, which take interleaved sequences of images and text and answer in text. Idefics3-8B pairs a SigLIP image encoder with Llama 3.1 8B Instruct and was built to improve OCR, document understanding and visual reasoning over Idefics2. Hugging Face trained it on open datasets it also publishes, including The Cauldron and Docmatix.
Idefics3 is released under Apache 2.0 although its language model is Meta's Llama 3.1 8B Instruct, and the card does not address the Llama license. SmolVLM, a separate entry, continues the architecture.
Openness
3 medium confidence- weights
- open(four safetensors shards on the Hub, ungated)
- data
- partial(fine-tuned on The Cauldron and Docmatix, which Hugging Face publishes
- code
- partial(fine-tuning examples only, an Idefics3 tutorial plus Idefics2 TRL and Trainer examples
- license
- Apache-2.0(OSI
- "We release the Idefics3 checkpoints under the Apache 2.0 license")+Llama-3.1-Community-License(carried from the Llama 3.1 8B Instruct base
- derivatives are licensed under it, with "Built with Llama" attribution and a 700-million-monthly-active-user bound on commercial use)
Idefics3's own checkpoints are released under Apache 2.0, but its language model is Llama 3.1, whose license carries over to the weights and bounds commercial use only for the very largest companies. Hugging Face publishes the main datasets it was fine-tuned on, but not the sets added for Idefics3 or its training code.
- https://arxiv.org/abs/2408.12637 recorded 2026-09-27
"trained efficiently, exclusively on open datasets"; "We release the model along with the datasets created for its training"
- https://cdn.jsdelivr.net/gh/meta-llama/llama-models@main/models/llama3_1/LICENSE recorded 2026-09-27
LLAMA 3.1 COMMUNITY LICENSE AGREEMENT: a model trained or fine-tuned from the Llama Materials must "include “Llama” at the beginning of any such AI model name"; licensees above "700 million monthly active users" must request a license from Meta
- https://huggingface.co/api/models/HuggingFaceM4/Idefics3-8B-Llama3?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for Idefics3-8B-Llama3: gated false, four safetensors shards, card license apache-2.0, datasets OBELICS, the_cauldron, Docmatix and WebSight
- https://huggingface.co/datasets/HuggingFaceM4/Docmatix/raw/main/README.md recorded 2026-09-27
Docmatix, MIT: "used for the fine-tuning of the vision-language model Idefics3"
- https://huggingface.co/datasets/HuggingFaceM4/the_cauldron/raw/main/README.md recorded 2026-09-27
The Cauldron: "a massive collection of 50 vision-language datasets (training sets only)", each governed by its own license
- https://huggingface.co/HuggingFaceM4/Idefics3-8B-Llama3/raw/main/README.md recorded 2026-09-27
"The model is built on top of two pre-trained models: google/siglip-so400m-patch14-384 and meta-llama/Meta-Llama-3.1-8B-Instruct"; "We release the Idefics3 checkpoints under the Apache 2.0 license"; "we have extended The Cauldron and added several datasets, including Docmatix. We will push soon these datasets to the same repo of The Cauldron (TODO)"; the card links a fine-tuning tutorial and Idefics2 TRL and Trainer examples, no training code
- https://ungh.cc/repos/merveenoyan/smol-vision/files/main recorded 2026-09-27
The smol-vision tree the card links holds Idefics_FT.ipynb, a fine-tuning notebook
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across Hugging Face's Idefics, Idefics2 and Idefics3 checkpoints. Almost all of it is Idefics3 and Idefics2 8B.
- https://huggingface.co/api/models?author=HuggingFaceM4&search=idefics&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=lastModified&expand[]=cardData&expand[]=gated recorded 2026-09-27
Fourteen Idefics checkpoints, two of them tiny-random test models, with 259,183 downloads in the trailing 30 days; Idefics3-8B-Llama3 127,897 and idefics2-8b 98,214
Capability
2 high confidenceIdefics3 reads whole documents and answers questions over several images at once, with a DocVQA score well above Idefics2's. Nothing on its card covers video, so it sits level with Moondream rather than with MiniCPM-V.
- https://huggingface.co/HuggingFaceM4/Idefics3-8B-Llama3/raw/main/README.md recorded 2026-09-27
"accepts arbitrary sequences of image and text inputs and produces text outputs"; "significantly enhancing capabilities around OCR, document understanding and visual reasoning"; Idefics3-8B row 46.6 / 58.4 / 55.9 / 87.7 / 74.9
Verified 2026-09-27