AI Potluck
Back to Gap Map Model components / Speech & audio

Qwen3-ASR

Alibaba Cloud
open weights / Overall score: 4.6(leading)

Alibaba's Qwen3-ASR speech recognition models (1.7B and 0.6B), built on the Qwen3 audio stack. They identify the language and transcribe 30 languages and 22 Chinese dialects, and a companion forced-alignment model predicts timestamps.

Openness

3 high confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
described(large-scale speech training data, not released)
code
partial(inference package and fine-tuning scripts, not the training recipe)
license
Apache-2.0(OSI)

Qwen3-ASR's weights and code are Apache-2.0, an OSI-approved license. The card mentions large-scale speech training data without releasing it, and the repository offers fine-tuning scripts rather than the recipe that produced the checkpoints, which keeps this at open weights.

Adoption

4 high confidence
4.0

Hugging Face downloads of the declared checkpoints, including the transformers-format copy the leaderboard evaluates. The qwen-asr package is marked as not a separate channel because it loads these same repositories.

Capability

5 medium confidence
5.0

Qwen3-ASR is the most accurate open recognizer on the English board and one of the broadest in language coverage; only proprietary services transcribe English with fewer errors.

Verified 2026-09-27