Granite Speech
IBMIBM's speech models: Granite Speech 4.1 pairs a speech encoder with a Granite language model for recognition and translation in English, French, German, Spanish, Portuguese and Japanese, and Granite Speech 5.0 TurboCTC is a compact 470M-parameter English recognizer built for speed.
The code repository carries no LICENSE file; its README states Apache 2.0.
Openness
3 medium confidence- weights
- open(safetensors and GGUF checkpoints on the Hub, ungated)
- data
- partial(third-party public corpora named in each card, plus synthetic sets that are not released)
- code
- partial(fine-tuning notebooks and inference
- license
- Apache-2.0(granite-speech-5.0-470m-turboctc and every 3.x and 4.x checkpoint)
- substitutable-variant
- granite-speech-5.0-470m-turboctc-nc(the same 470M architecture trained with GigaSpeech and SPGISpeech added, under CC-BY-NC-SA-4.0
IBM's current Granite Speech 5.0 model and every earlier checkpoint are Apache-2.0 and download freely. The training data are public corpora plus synthetic sets IBM has not released, and only fine-tuning notebooks are published. A non-commercial 5.0 variant exists, but it is the same model trained on more data and IBM points commercial users to the Apache-2.0 one.
- https://huggingface.co/api/models?author=ibm-granite&search=speech&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Family listing: eleven ibm-granite speech repositories tagged "license:apache-2.0" and granite-speech-5.0-470m-turboctc-nc tagged "license:cc-by-nc-sa-4.0", all ungated.
- https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc-nc/raw/main/README.md recorded 2026-09-27
Card: "This model is intended for research and non-commercial use-only. See [Granite Speech 5.0 Turbo CTC](https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc) if interested in commercial use."; "It was trained on approximately 75,000 hours of English audio from public corpora".
- https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc/raw/main/README.md recorded 2026-09-27
Card: "Granite Speech 5.0 TurboCTC is a compact 470 million parameter English ASR model"; "It was trained on approximately 60,000 hours of English audio from public corpora"; "Our training data is entirely comprised of publicly available datasets or of synthetic data generated from public corpora".
- https://raw.githubusercontent.com/ibm-granite/granite-speech-models/main/notebooks/finetune_granite_speech_5_turboctc.ipynb recorded 2026-09-27
Fine-tuning notebook for Granite Speech 5.0 TurboCTC in the code repository.
- https://raw.githubusercontent.com/ibm-granite/granite-speech-models/main/README.md recorded 2026-09-27
README: "**License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)".
Adoption
3 high confidenceHugging Face downloads summed over the five declared checkpoints, led by the multilingual 4.1 models.
- https://huggingface.co/api/models?author=ibm-granite&search=speech&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: granite-speech-4.1-2b 154,165, 4.1-2b-plus 112,757, 4.1-2b-nar 82,910, 3.3-2b 71,840, 5.0-470m-turboctc 52,451.
Capability
4 medium confidenceAmong the most accurate open recognizers on the English board, a step below Qwen3-ASR, and the 4.1 models add speech translation across six languages.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV rows "ibm-granite/granite-speech-4.1-2b,4.615" and "ibm-granite/granite-speech-5.0-470m-turboctc,5.04375".
- https://huggingface.co/ibm-granite/granite-speech-4.1-2b/raw/main/README.md recorded 2026-09-27
Card: "specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)".
Verified 2026-09-27