AI Potluck
Back to Gap Map Model components / Speech & audio

Canary

NVIDIA
open weights / Overall score: 3.2

NVIDIA's multitask speech recognition and translation models. Canary 1B v2 transcribes 25 European languages and translates between English and 24 of them; Canary-Qwen-2.5B pairs the Canary encoder with a Qwen language model for English transcription.

The current checkpoints (1B v2, the Flash models and Canary-Qwen-2.5B) are CC-BY-4.0, and the Canary-Qwen checkpoint pairs the Canary encoder with a Qwen language model.

Openness

3 high confidence
3.0
weights
open(.nemo and safetensors checkpoints, ungated)
data
partial(Granary is released
code
open(NeMo speech_to_text_aed.py training script named in the card)
license
CC-BY-4.0(canary-1b-v2, the current release
superseded-release
canary-1b(2024, still published under CC-BY-NC-4.0

Canary's current release, Canary 1B v2 (August 2025), downloads freely under CC-BY-4.0, an attribution-only license, and its card names the NeMo training script. Its Granary training data is public but NeMo ASR Set 3.0 is in-house, so the model cannot be rebuilt from what is released. The Flash and Canary-Qwen checkpoints carry the same license. The 2024 canary-1b is still published under a non-commercial license, but it is a superseded release and does not set the score.

Adoption

2 medium confidence
2.0

Hugging Face downloads summed over the four current checkpoints, with Canary 1B v2 and Canary-Qwen-2.5B carrying most of them.

Capability

4 medium confidence
4.0

The Canary line is split between accuracy and breadth: Canary-Qwen sits near the top of the English board, while Canary 1B v2 trades some English accuracy for 25-language transcription and translation.

Verified 2026-09-27