AI Potluck
Back to Gap Map Model components / Multimodal models

Molmo

Allen Institute for AI
open source / Overall score: 3.0

Molmo is Ai2's family of open vision-language models, released with their training data and code. Molmo2 understands single images, several images and video, and answers by pointing at objects or tracking them across frames as well as in text. Its checkpoints pair Qwen3 or OLMo language models with a SigLIP 2 vision encoder, and every dataset was built without distilling from closed models.

Ai2's robot-action MolmoAct models are a separate line. The model cards note that some of the third-party datasets Molmo was trained on are limited to non-commercial research use.

Openness

5 high confidence
5.0
weights
open(safetensors on the Hub, ungated)
data
open(the Molmo2 Data collection and PixMo sets are published ungated, mostly ODC-BY
code
open(pre-training, SFT and long-context SFT launch scripts in allenai/molmo2)
license
Apache-2.0(OSI

Ai2 publishes Molmo2's weights under Apache 2.0 together with the datasets it was trained on and the full training code, so the model can be rebuilt from what ships. The datasets were built without distilling from closed models.

Adoption

3 high confidence
3.0

Hugging Face downloads over the trailing 30 days, summed across Ai2's Molmo and Molmo2 checkpoints. MolmoAct, the embodied-reasoning Molmo2-ER, MolmoPoint, MolmoWeb and the other Molmo-named lines are excluded.

Capability

3 high confidence
3.0

Molmo2 understands images, image sets and video and can point at or track what it describes, where it beats much larger models. It reports no GUI operation or audio input, so it sits level with MiniCPM-V.

Verified 2026-09-26