Molmo
Allen Institute for AIMolmo is Ai2's family of open vision-language models, released with their training data and code. Molmo2 understands single images, several images and video, and answers by pointing at objects or tracking them across frames as well as in text. Its checkpoints pair Qwen3 or OLMo language models with a SigLIP 2 vision encoder, and every dataset was built without distilling from closed models.
Ai2's robot-action MolmoAct models are a separate line. The model cards note that some of the third-party datasets Molmo was trained on are limited to non-commercial research use.
Openness
5 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- open(the Molmo2 Data collection and PixMo sets are published ungated, mostly ODC-BY
- code
- open(pre-training, SFT and long-context SFT launch scripts in allenai/molmo2)
- license
- Apache-2.0(OSI
Ai2 publishes Molmo2's weights under Apache 2.0 together with the datasets it was trained on and the full training code, so the model can be rebuilt from what ships. The datasets were built without distilling from closed models.
- https://arxiv.org/pdf/2601.10611 recorded 2026-09-26
"we release our training data, model weights, and training code ... all our data is constructed without distilling from proprietary models"
- https://huggingface.co/api/collections/allenai/molmo2-data recorded 2026-09-26
The "Molmo2 Data" collection: Molmo2-Cap, VideoCapQA, VideoSubtitleQA, AskModelAnything, VideoPoint, VideoTrack, MultiImageQA, SynMultiImageQA and MultiImagePoint
- https://huggingface.co/api/datasets/allenai/Molmo2-Cap recorded 2026-09-26
Molmo2-Cap: public, ungated, odc-by
- https://huggingface.co/api/models/allenai/Molmo2-8B?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for Molmo2-8B: gated false, eight safetensors shards, card license apache-2.0, nine Molmo2 training datasets listed
- https://raw.githubusercontent.com/allenai/molmo2/HEAD/launch_scripts/sft.py recorded 2026-09-26
The multitask SFT launch script in allenai/molmo2
- https://raw.githubusercontent.com/allenai/molmo2/HEAD/LICENSE recorded 2026-09-26
The repository LICENSE is the Apache License, Version 2.0
- https://raw.githubusercontent.com/allenai/molmo2/HEAD/scripts/download_datasets.py recorded 2026-09-27
The script that downloads the training datasets
Adoption
3 high confidenceHugging Face downloads over the trailing 30 days, summed across Ai2's Molmo and Molmo2 checkpoints. MolmoAct, the embodied-reasoning Molmo2-ER, MolmoPoint, MolmoWeb and the other Molmo-named lines are excluded.
- https://huggingface.co/api/models?author=allenai&search=Molmo&limit=100 recorded 2026-09-26
Eight Molmo and Molmo2 checkpoints with 245,590 downloads in the trailing 30 days, Molmo2-ER's 7,104 left out; Molmo2-4B 102,438 and Molmo2-8B 71,158
Capability
3 high confidenceMolmo2 understands images, image sets and video and can point at or track what it describes, where it beats much larger models. It reports no GUI operation or audio input, so it sits level with MiniCPM-V.
- https://arxiv.org/abs/2601.10611 recorded 2026-09-26
Abstract: "38.4 vs 20.0 F1 on video pointing and 56.2 vs 41.1 J&F on video tracking" against Gemini 3 Pro
- https://arxiv.org/pdf/2601.10611 recorded 2026-09-26
Video-MME 69.9 and MMMU val 53.0 for Molmo2-8B against 71.4 and 69.6 for Qwen3-VL-8B; maximum 128 frames, 384 for long-context training
- https://huggingface.co/allenai/Molmo2-8B/raw/main/README.md recorded 2026-09-26
"support image, video and multi-image understanding and grounding"; Pointing, Tracking and Multi-Image Point QA examples
Verified 2026-09-26