Canary
NVIDIANVIDIA's multitask speech recognition and translation models. Canary 1B v2 transcribes 25 European languages and translates between English and 24 of them; Canary-Qwen-2.5B pairs the Canary encoder with a Qwen language model for English transcription.
The current checkpoints (1B v2, the Flash models and Canary-Qwen-2.5B) are CC-BY-4.0, and the Canary-Qwen checkpoint pairs the Canary encoder with a Qwen language model.
Openness
3 high confidence- weights
- open(.nemo and safetensors checkpoints, ungated)
- data
- partial(Granary is released
- code
- open(NeMo speech_to_text_aed.py training script named in the card)
- license
- CC-BY-4.0(canary-1b-v2, the current release
- superseded-release
- canary-1b(2024, still published under CC-BY-NC-4.0
Canary's current release, Canary 1B v2 (August 2025), downloads freely under CC-BY-4.0, an attribution-only license, and its card names the NeMo training script. Its Granary training data is public but NeMo ASR Set 3.0 is in-house, so the model cannot be rebuilt from what is released. The Flash and Canary-Qwen checkpoints carry the same license. The 2024 canary-1b is still published under a non-commercial license, but it is a superseded release and does not set the score.
- https://huggingface.co/api/models?author=nvidia&search=canary recorded 2026-09-27
Family listing with createdAt: canary-1b-v2 2025-08-04, canary-qwen-2.5b 2025-06-26, canary-180m-flash 2025-03-11 and canary-1b-flash 2025-03-07, all tagged "license:cc-by-4.0"; canary-1b 2024-02-07, tagged "license:cc-by-nc-4.0".
- https://huggingface.co/api/models/nvidia/canary-1b recorded 2026-09-27
nvidia/canary-1b is public, "createdAt":"2024-02-07T17:20:55.000Z", and tagged "license:cc-by-nc-4.0".
- https://huggingface.co/api/models/nvidia/canary-1b-v2 recorded 2026-09-27
"license:cc-by-4.0" tag, "gated":false, "createdAt":"2025-08-04T13:34:41.000Z", and canary-1b-v2.nemo and model.safetensors in the file list.
- https://huggingface.co/nvidia/canary-1b-v2/raw/main/README.md recorded 2026-09-27
Card: "combining Nvidia's newly published [Granary](https://huggingface.co/datasets/nvidia/Granary) and in-house dataset NeMo ASR Set 3.0"; "Training script: [speech\_to\_text\_aed.py]"; Release Date "Huggingface [08/14/2025]"; "GOVERNING TERMS: Use of this model is governed by the [CC-BY-4.0]".
Adoption
2 medium confidenceHugging Face downloads summed over the four current checkpoints, with Canary 1B v2 and Canary-Qwen-2.5B carrying most of them.
- https://huggingface.co/api/models?author=nvidia&search=canary recorded 2026-09-27
Trailing-30-day downloads: canary-1b-v2 56,469, canary-qwen-2.5b 35,182, canary-1b-flash 3,937, canary-180m-flash 2,955.
Capability
4 medium confidenceThe Canary line is split between accuracy and breadth: Canary-Qwen sits near the top of the English board, while Canary 1B v2 trades some English accuracy for 25-language transcription and translation.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV rows "nvidia/canary-qwen-2.5b,4.4275" and "nvidia/canary-1b-v2,5.70625".
- https://huggingface.co/nvidia/canary-1b-v2/raw/main/README.md recorded 2026-09-27
"Speech Transcription (ASR) for 25 languages"; "Speech Translation (AST) from English → 24 languages".
Verified 2026-09-27