AI Potluck
Back to Gap Map Model components / Speech & audio

NVIDIA NeMo Speech

NVIDIA
open source / Overall score: 4.2(strong)

NVIDIA's PyTorch toolkit for training, fine-tuning and serving speech models: recognition, synthesis, speaker tasks and speech LLMs. It is the framework behind NVIDIA's Parakeet, Canary and Magpie TTS models and installs as nemo-toolkit.

The repository moved from NVIDIA/NeMo to NVIDIA-NeMo/Speech, and NVIDIA/NeMo resolves to it.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public
core features withheld
no

NeMo Speech is Apache-2.0 and the published repository is the whole toolkit. NVIDIA sells hosted inference as its own NIM product, which does not withhold anything from the open source.

Adoption

3 high confidence
3.0

PyPI downloads of nemo-toolkit, the package the README installs with its ASR and TTS extras.

Capability

5 medium confidence
5.0

The only engine here that trains frontier-class recognizers as well as running them, across both recognition and synthesis.

  • https://raw.githubusercontent.com/NVIDIA-NeMo/Speech/HEAD/README.md recorded 2026-09-27

    "Recognition (ASR), Text to Speech (TTS), and Speech LLMs. It is designed to help you efficiently create, customize, and"; "[Nemotron-3.5-ASR-Streaming-0.6B](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) has been released with 40 languages supported".

Verified 2026-09-27