AI Potluck
Back to Gap Map Model components / Multimodal models

Qwen-VL

Alibaba Cloud
open weights / Overall score: 4.3(strong)

Qwen-VL is Alibaba's vision-language model line, from Qwen-VL and Qwen2-VL through Qwen2.5-VL to Qwen3-VL, released in dense and mixture-of-experts sizes with Instruct and Thinking editions. Qwen3-VL reads images, documents in 32 OCR languages and hours-long video in a 256K context, grounds objects in 2D and 3D, and operates desktop and phone interfaces. Its successor is the natively multimodal Qwen3.5 family.

Scored on Qwen3-VL. Earlier generations are still downloaded, and some of their larger or smaller sizes carry the Qwen or Qwen Research license rather than Apache 2.0.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(the technical report says the SFT set was curated from open-source datasets and web resources and the RL set of about 30K queries from open-source and proprietary sources
code
partial(a supervised fine-tuning framework for users (qwen-vl-finetune)
license
Apache-2.0(OSI

Every Qwen3-VL checkpoint is released under Apache 2.0, and the repository includes a fine-tuning framework. The training data is described in the technical report but not published, and Alibaba's own training recipe is not released, so the model cannot be rebuilt from what ships.

Adoption

5 high confidence
5.0

Hugging Face downloads over the trailing 30 days, summed across Alibaba's own Qwen-VL, Qwen2-VL, Qwen2.5-VL, Qwen3-VL and QVQ checkpoints, including its own quantized builds. The Qwen3-VL embedding and reranker models are a separate line and are excluded.

Capability

4 high confidence
4.0

Qwen3-VL operates desktop and phone interfaces from screenshots and reports OSWorld and AndroidWorld results, on top of long-video and document understanding. It is level with UI-TARS and short of Qwen-Omni because it takes no audio input.

Verified 2026-09-26