AI Potluck
Back to Gap Map Model components / Multimodal models

Qwen-Omni

Alibaba Cloud
closed / Overall score: 4.7(leading)

Qwen-Omni is Alibaba's omni-modal model line, which takes text, images, audio and video as input and answers in text or in streaming natural speech. Its current release, Qwen3.5-Omni, is served only through Alibaba Cloud's API. The earlier Qwen3-Omni, a mixture-of-experts Thinker-Talker model with Instruct, Thinking and audio Captioner checkpoints, and Qwen2.5-Omni before it remain downloadable.

The current release, Qwen3.5-Omni, is hosted-only, while Qwen3-Omni and Qwen2.5-Omni remain downloadable under Apache 2.0. It scores 1/closed on the current release, and the older open line trails it. A separate availability attribute that would record that gap is tracked in #727.

Openness

1 high confidence
1.0
weights
closed(Qwen3.5-Omni is served only through Alibaba Cloud Model Studio
data
closed(nothing is published for Qwen3.5-Omni)
code
closed(no code is published for Qwen3.5-Omni)
license
Proprietary(Qwen3.5-Omni is available only under Model Studio API terms)
prior-release
distributed(Qwen3-Omni (and Qwen2.5-Omni): safetensors on the Hub, ungated, under Apache-2.0, with inference cookbooks and serving code only and no post-training data)

Qwen3.5-Omni, the current release, can only be used through Alibaba Cloud's API, and no weights have been published for it; Qwen3.8-Omni-Flash is hosted only too. The previous generation, Qwen3-Omni, remains downloadable under Apache 2.0. The open line trails the vendor's closed release. A separate availability attribute that would record that gap is tracked in #727.

Adoption

4 high confidence
4.0

Hugging Face downloads over the trailing 30 days, summed across Alibaba's own Qwen3-Omni and Qwen2.5-Omni checkpoints. Use of the hosted Qwen3.5-Omni API is not counted.

Capability

5 high confidence
5.0

Qwen3-Omni understands audio and video natively alongside images and text and answers in streaming speech as well as text, the widest range of inputs and outputs in this category. Its vision scores sit close to Qwen2.5-VL-72B on the card's own comparison, so the breadth does not come at a large cost in image understanding.

Verified 2026-09-26