Qwen-VL
Alibaba CloudQwen-VL is Alibaba's vision-language model line, from Qwen-VL and Qwen2-VL through Qwen2.5-VL to Qwen3-VL, released in dense and mixture-of-experts sizes with Instruct and Thinking editions. Qwen3-VL reads images, documents in 32 OCR languages and hours-long video in a 256K context, grounds objects in 2D and 3D, and operates desktop and phone interfaces. Its successor is the natively multimodal Qwen3.5 family.
Scored on Qwen3-VL. Earlier generations are still downloaded, and some of their larger or smaller sizes carry the Qwen or Qwen Research license rather than Apache 2.0.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- described(the technical report says the SFT set was curated from open-source datasets and web resources and the RL set of about 30K queries from open-source and proprietary sources
- code
- partial(a supervised fine-tuning framework for users (qwen-vl-finetune)
- license
- Apache-2.0(OSI
Every Qwen3-VL checkpoint is released under Apache 2.0, and the repository includes a fine-tuning framework. The training data is described in the technical report but not published, and Alibaba's own training recipe is not released, so the model cannot be rebuilt from what ships.
- https://arxiv.org/html/2511.21631 recorded 2026-09-26
Qwen3-VL technical report: SFT data curated "from open-source datasets and web resources"; RL data "from both open-source and proprietary sources", about 30K queries; no release link
- https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-VL@main/LICENSE recorded 2026-09-26
The repository LICENSE is the Apache License, Version 2.0
- https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-VL@main/qwen-vl-finetune/README.md recorded 2026-09-26
"This repository provides a training framework for Qwen VL models": a supervised fine-tuning trainer, argument and data-packing tools
- https://huggingface.co/api/models?author=Qwen&search=VL&limit=100 recorded 2026-09-26
Qwen-owned VL listing: every Qwen3-VL checkpoint is tagged license:apache-2.0; some Qwen2-VL and Qwen2.5-VL 72B checkpoints are tagged license:other
- https://huggingface.co/api/models/Qwen/Qwen3-VL-8B-Instruct?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for Qwen3-VL-8B-Instruct: gated false, four safetensors shards, card license apache-2.0
Adoption
5 high confidenceHugging Face downloads over the trailing 30 days, summed across Alibaba's own Qwen-VL, Qwen2-VL, Qwen2.5-VL, Qwen3-VL and QVQ checkpoints, including its own quantized builds. The Qwen3-VL embedding and reranker models are a separate line and are excluded.
- https://huggingface.co/api/models?author=Qwen&search=QVQ&limit=100 recorded 2026-09-26
QVQ-72B-Preview, 603 downloads in the trailing 30 days
- https://huggingface.co/api/models?author=Qwen&search=VL&limit=100 recorded 2026-09-26
62 Qwen-owned vision-language chat checkpoints (the four Qwen3-VL-Embedding and Reranker repositories excluded); trailing-30-day downloads sum to 46,805,754, Qwen3-VL-8B-Instruct alone 18,555,061
Capability
4 high confidenceQwen3-VL operates desktop and phone interfaces from screenshots and reports OSWorld and AndroidWorld results, on top of long-video and document understanding. It is level with UI-TARS and short of Qwen-Omni because it takes no audio input.
- https://arxiv.org/html/2511.21631 recorded 2026-09-26
"Qwen3-VL 32B scores 41 on OSWorld and 63.7 on AndroidWorld"; ScreenSpot, ScreenSpot Pro and OSWorldG among the evaluated GUI benchmarks
- https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct/raw/main/README.md recorded 2026-09-26
"Visual Agent: Operates PC/mobile GUIs"; "Native 256K context, expandable to 1M; handles books and hours-long video"; OCR in 32 languages; 2D and 3D grounding
Verified 2026-09-26