AI Potluck
Back to Gap Map Model components / Multimodal models

VideoLLaMA

Alibaba DAMO Academy
open weights / Overall score: 4.1(strong)

VideoLLaMA is Alibaba DAMO Academy's video-language model line, which answers questions about video and images in text. VideoLLaMA3 builds on Qwen2.5 language models in 2B and 7B sizes, samples video at a configurable frame rate and is trained in four stages from vision-encoder adaptation to video-centric fine-tuning. VideoLLaMA2.1-AV adds audio-visual input.

The repository's README calls the service a research preview for non-commercial use, while the LICENSE file and every checkpoint carry Apache 2.0; the entry reads the license, not the README.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
partial(the VL3-Syn7M re-captioned image set is released
code
open(training scripts for every stage (scripts/train) and an evaluation harness)
license
Apache-2.0(OSI

VideoLLaMA3's weights and code are Apache 2.0, and the repository includes training scripts for every stage. DAMO releases its re-captioned image set, but the full training mixture is assembled from third-party datasets rather than published as one release.

Adoption

2 high confidence
2.0

Hugging Face downloads over the trailing 30 days, summed across DAMO's VideoLLaMA2, 2.1 and 3 checkpoints. The older Video-LLaMA repositories record no downloads.

Capability

5 medium confidence
5.0

VideoLLaMA2.1-AV understands a video's soundtrack together with its frames, and VideoLLaMA3 reads long frame sequences and still images. That audio-visual input puts the line level with Qwen-Omni, although it answers only in text and its audio-visual model is an older, smaller release.

Verified 2026-09-27