Qwen-Omni
Alibaba CloudQwen-Omni is Alibaba's omni-modal model line, which takes text, images, audio and video as input and answers in text or in streaming natural speech. Its current release, Qwen3.5-Omni, is served only through Alibaba Cloud's API. The earlier Qwen3-Omni, a mixture-of-experts Thinker-Talker model with Instruct, Thinking and audio Captioner checkpoints, and Qwen2.5-Omni before it remain downloadable.
The current release, Qwen3.5-Omni, is hosted-only, while Qwen3-Omni and Qwen2.5-Omni remain downloadable under Apache 2.0. It scores 1/closed on the current release, and the older open line trails it. A separate availability attribute that would record that gap is tracked in #727.
Openness
1 high confidence- weights
- closed(Qwen3.5-Omni is served only through Alibaba Cloud Model Studio
- data
- closed(nothing is published for Qwen3.5-Omni)
- code
- closed(no code is published for Qwen3.5-Omni)
- license
- Proprietary(Qwen3.5-Omni is available only under Model Studio API terms)
- prior-release
- distributed(Qwen3-Omni (and Qwen2.5-Omni): safetensors on the Hub, ungated, under Apache-2.0, with inference cookbooks and serving code only and no post-training data)
Qwen3.5-Omni, the current release, can only be used through Alibaba Cloud's API, and no weights have been published for it; Qwen3.8-Omni-Flash is hosted only too. The previous generation, Qwen3-Omni, remains downloadable under Apache 2.0. The open line trails the vendor's closed release. A separate availability attribute that would record that gap is tracked in #727.
- https://api.github.com/search/repositories?q=org:QwenLM+omni recorded 2026-09-27
GitHub search of the QwenLM organization for omni: four repositories, Qwen3-Omni, Qwen2.5-Omni, Qwen-Live-Harness and Omnilingua-Bench; no Qwen3.5-Omni code or data repository.
- https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-Omni@main/LICENSE recorded 2026-09-26
The repository LICENSE is the Apache License, Version 2.0, with no added field-of-use or revenue clause
- https://cdn.jsdelivr.net/gh/QwenLM/Qwen3-Omni@main/README.md recorded 2026-09-26
Repository README: model downloads, inference and vLLM serving instructions and cookbooks; no training scripts and no post-training dataset links
- https://huggingface.co/api/models?author=Qwen&search=Qwen3.5-Omni&limit=100 recorded 2026-09-27
An empty list: no Qwen3.5-Omni weights are published under the Qwen organization on the Hub
- https://huggingface.co/api/models/Qwen/Qwen3-Omni-30B-A3B-Instruct?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=siblings recorded 2026-09-26
Hub metadata for Qwen3-Omni-30B-A3B-Instruct: gated false, safetensors weight shards, card license other with license_name apache-2.0
- https://www.alibabacloud.com/help/en/model-studio/qwen-omni recorded 2026-09-27
Model Studio Qwen-Omni documentation (last updated Sep 18, 2026): "Choose Qwen3.8-Omni-Flash for text analysis and Qwen3.5-Omni for speech output", called through the DashScope API as qwen3.5-omni-plus and qwen3.8-omni-flash with an API key; no download.
Adoption
4 high confidenceHugging Face downloads over the trailing 30 days, summed across Alibaba's own Qwen3-Omni and Qwen2.5-Omni checkpoints. Use of the hosted Qwen3.5-Omni API is not counted.
- https://huggingface.co/api/models?author=Qwen&search=Omni&limit=100 recorded 2026-09-26
Seven Qwen-owned Omni checkpoints; their trailing-30-day downloads sum to 1,548,543, led by Qwen3-Omni-30B-A3B-Instruct at 612,375
Capability
5 high confidenceQwen3-Omni understands audio and video natively alongside images and text and answers in streaming speech as well as text, the widest range of inputs and outputs in this category. Its vision scores sit close to Qwen2.5-VL-72B on the card's own comparison, so the breadth does not come at a large cost in image understanding.
- https://arxiv.org/abs/2509.17765 recorded 2026-09-27
Abstract: across 36 audio and audio-visual benchmarks Qwen3-Omni reaches open-source state of the art on 32, with a theoretical first-packet latency of 234 ms
- https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/raw/main/README.md recorded 2026-09-27
"It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech." Vision table: MMMU_val 69.1, MathVista_mini 75.9, Video-MME 70.5
Verified 2026-09-26