AI Potluck
Back to Gap Map Model components / Multimodal models

LLaVA

LMMs-Lab
open source / Overall score: 3.3

LLaVA is the open vision-language model line that began with visual instruction tuning and continued through LLaVA-NeXT and LLaVA-OneVision. The current LLaVA-OneVision-2 is an 8B model on a Qwen3 base that handles single images, multi-image sets and long video, and is released with its data, vision encoders, training code and logs. Hugging Face maintains transformers conversions of the earlier releases.

Scored on LLaVA-OneVision-2. Its weights sit under the lmms-lab-encoder account and its data under mvp-lab, and older checkpoints built on Llama 2, Llama 3 or Vicuna carry those bases' licenses.

Openness

5 high confidence
5.0
weights
open(safetensors on the Hub, ungated)
data
open(LLaVA-OneVision-2-Data is published ungated under Apache 2.0, with the OneVision-1.5 mid-training and instruct sets carried forward)
code
open(training scripts and configs in EvolvingLMMs-Lab/LLaVA-OneVision-1.5 (examples/llava_onevision2))
license
Apache-2.0(OSI

LLaVA-OneVision-2 is released under Apache 2.0 with its training data, vision encoders, training code and logs, so the model can be reproduced from what ships.

Adoption

4 medium confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across the lmms-lab and lmms-lab-encoder checkpoints and Hugging Face's llava-hf conversions. Nearly all of it is the older LLaVA-1.5 and 1.6 conversions rather than the current OneVision releases.

Capability

3 high confidence
3.0

LLaVA-OneVision-2 reads single images, image sets and long video, and its predecessor already edged past Qwen2.5-VL-7B on documents and charts. It reports no GUI operation or audio input, so it sits level with MiniCPM-V.

Verified 2026-09-27