AI Potluck
Back to Gap Map Model components / Speech & audio

Granite Speech

IBM
open weights / Overall score: 3.6

IBM's speech models: Granite Speech 4.1 pairs a speech encoder with a Granite language model for recognition and translation in English, French, German, Spanish, Portuguese and Japanese, and Granite Speech 5.0 TurboCTC is a compact 470M-parameter English recognizer built for speed.

The code repository carries no LICENSE file; its README states Apache 2.0.

Openness

3 medium confidence
3.0
weights
open(safetensors and GGUF checkpoints on the Hub, ungated)
data
partial(third-party public corpora named in each card, plus synthetic sets that are not released)
code
partial(fine-tuning notebooks and inference
license
Apache-2.0(granite-speech-5.0-470m-turboctc and every 3.x and 4.x checkpoint)
substitutable-variant
granite-speech-5.0-470m-turboctc-nc(the same 470M architecture trained with GigaSpeech and SPGISpeech added, under CC-BY-NC-SA-4.0

IBM's current Granite Speech 5.0 model and every earlier checkpoint are Apache-2.0 and download freely. The training data are public corpora plus synthetic sets IBM has not released, and only fine-tuning notebooks are published. A non-commercial 5.0 variant exists, but it is the same model trained on more data and IBM points commercial users to the Apache-2.0 one.

Adoption

3 high confidence
3.0

Hugging Face downloads summed over the five declared checkpoints, led by the multilingual 4.1 models.

Capability

4 medium confidence
4.0

Among the most accurate open recognizers on the English board, a step below Qwen3-ASR, and the 4.1 models add speech translation across six languages.

Verified 2026-09-27