Qwen3-ASR
Alibaba CloudAlibaba's Qwen3-ASR speech recognition models (1.7B and 0.6B), built on the Qwen3 audio stack. They identify the language and transcribe 30 languages and 22 Chinese dialects, and a companion forced-alignment model predicts timestamps.
Openness
3 high confidence- weights
- open(safetensors on the Hub, ungated)
- data
- described(large-scale speech training data, not released)
- code
- partial(inference package and fine-tuning scripts, not the training recipe)
- license
- Apache-2.0(OSI)
Qwen3-ASR's weights and code are Apache-2.0, an OSI-approved license. The card mentions large-scale speech training data without releasing it, and the repository offers fine-tuning scripts rather than the recipe that produced the checkpoints, which keeps this at open weights.
- https://huggingface.co/api/models/Qwen/Qwen3-ASR-1.7B recorded 2026-09-27
"license:apache-2.0" tag, "gated":false, with two safetensors shards.
- https://raw.githubusercontent.com/QwenLM/Qwen3-ASR/HEAD/LICENSE recorded 2026-09-27
LICENSE body: "Apache License" "Version 2.0, January 2004".
- https://raw.githubusercontent.com/QwenLM/Qwen3-ASR/HEAD/README.md recorded 2026-09-27
"Both leverage large-scale speech training data"; "Please refer to [Qwen3-ASR-Finetuning](finetuning/) for detailed instructions on fine-tuning Qwen3-ASR."
Adoption
4 high confidenceHugging Face downloads of the declared checkpoints, including the transformers-format copy the leaderboard evaluates. The qwen-asr package is marked as not a separate channel because it loads these same repositories.
- https://huggingface.co/api/models?author=Qwen&search=Qwen3-ASR recorded 2026-09-27
Qwen/Qwen3-ASR-0.6B 553,397 and Qwen/Qwen3-ASR-1.7B-hf 320,009 downloads in the trailing 30 days.
- https://huggingface.co/api/models/Qwen/Qwen3-ASR-1.7B recorded 2026-09-27
"downloads":1924880 for Qwen/Qwen3-ASR-1.7B.
Capability
5 medium confidenceQwen3-ASR is the most accurate open recognizer on the English board and one of the broadest in language coverage; only proprietary services transcribe English with fewer errors.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV row "Qwen/Qwen3-ASR-1.7B-hf,4.31125,819.96,apache-2.0" (size 2.04B, 52 languages); every row with a lower average carries "Proprietary".
- https://raw.githubusercontent.com/QwenLM/Qwen3-ASR/HEAD/README.md recorded 2026-09-27
"support language identification and ASR for 52 languages and dialects".
Verified 2026-09-27