VideoLLaMA
Alibaba DAMO AcademyVideoLLaMA is Alibaba DAMO Academy's video-language model line, which answers questions about video and images in text. VideoLLaMA3 builds on Qwen2.5 language models in 2B and 7B sizes, samples video at a configurable frame rate and is trained in four stages from vision-encoder adaptation to video-centric fine-tuning. VideoLLaMA2.1-AV adds audio-visual input.
The repository's README calls the service a research preview for non-commercial use, while the LICENSE file and every checkpoint carry Apache 2.0; the entry reads the license, not the README.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- partial(the VL3-Syn7M re-captioned image set is released
- code
- open(training scripts for every stage (scripts/train) and an evaluation harness)
- license
- Apache-2.0(OSI
VideoLLaMA3's weights and code are Apache 2.0, and the repository includes training scripts for every stage. DAMO releases its re-captioned image set, but the full training mixture is assembled from third-party datasets rather than published as one release.
- https://cdn.jsdelivr.net/gh/DAMO-NLP-SG/VideoLLaMA3@main/LICENSE recorded 2026-09-27
The repository LICENSE is the Apache License, Version 2.0
- https://cdn.jsdelivr.net/gh/DAMO-NLP-SG/VideoLLaMA3@main/README.md recorded 2026-09-27
"We provide some templates in scripts/train for all stages"; License section: Apache 2.0, and "The service is a research preview intended for non-commercial use ONLY"
- https://cdn.jsdelivr.net/gh/DAMO-NLP-SG/VideoLLaMA3@main/scripts/train/stage1_2b.sh recorded 2026-09-27
The stage-1 training launch script for the 2B model
- https://huggingface.co/api/datasets/DAMO-NLP-SG/VL3-Syn7M recorded 2026-09-27
VL3-Syn7M: public, ungated, apache-2.0, "the re-captioned data we used during the training of VideoLLaMA3"
- https://huggingface.co/api/models/DAMO-NLP-SG/VideoLLaMA3-7B?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-27
Hub metadata for VideoLLaMA3-7B: gated false, four safetensors shards, license apache-2.0, trained on LLaVA-OneVision-Data, pixmo-docs, Docmatix, LLaVA-Video-178K and ShareGPT4Video
Adoption
2 high confidenceHugging Face downloads over the trailing 30 days, summed across DAMO's VideoLLaMA2, 2.1 and 3 checkpoints. The older Video-LLaMA repositories record no downloads.
- https://huggingface.co/api/models?author=DAMO-NLP-SG&search=VideoLLaMA&limit=100 recorded 2026-09-27
17 VideoLLaMA2, 2.1 and 3 checkpoints with 11,148 downloads in the trailing 30 days; VideoLLaMA3-7B 3,912 and VideoLLaMA3-2B 3,799
Capability
5 medium confidenceVideoLLaMA2.1-AV understands a video's soundtrack together with its frames, and VideoLLaMA3 reads long frame sequences and still images. That audio-visual input puts the line level with Qwen-Omni, although it answers only in text and its audio-visual model is an older, smaller release.
- https://arxiv.org/abs/2406.07476 recorded 2026-09-27
VideoLLaMA 2 integrates "an Audio Branch into the model through joint training" and improves on audio-only and audio-video question answering
- https://cdn.jsdelivr.net/gh/DAMO-NLP-SG/VideoLLaMA3@main/README.md recorded 2026-09-27
Example video input with fps 1 and max_frames 180; "VideoLLaMA3-7B is the best 7B-sized model on VideoMME leaderboard" as of Jan 24
- https://huggingface.co/DAMO-NLP-SG/VideoLLaMA2.1-7B-AV/raw/main/README.md recorded 2026-09-27
Tasks listed: Audio-visual Question Answering and Audio Question Answering; "Release checkpoints of VideoLLaMA2.1-7B-AV"
Verified 2026-09-27