AI Potluck
Back to Gap Map Model components / Speech & audio

Whisper

OpenAI
open weights / Overall score: 3.4

OpenAI's open-weight speech recognition and translation model family, from tiny to large-v3 and the faster large-v3-turbo. It transcribes about 99 languages and translates speech into English.

The large-v3 checkpoints carry Apache-2.0 on the Hub and large-v3-turbo carries MIT, the same license as the openai/whisper code. OpenAI's hosted transcription models are a separate product, openai-speech.

Openness

3 high confidence
3.0
weights
open(safetensors checkpoints on the Hub, ungated)
data
described(1M hours weakly labeled plus 4M hours pseudo-labeled audio, not released)
code
partial(inference and decoding code only)
license
MIT(large-v3-turbo weights and the code)+Apache-2.0(large-v3 and earlier weights)

Whisper's weights download without a gate under MIT and Apache-2.0, both OSI-approved. OpenAI describes the five million hours of audio it trained on but has not released them, and the repository holds inference code rather than a training pipeline, so this is open weights rather than a reproducible model.

Adoption

4 high confidence
4.0

Read from PyPI, the route that leads for a product declaring a package: installs of openai-whisper, the reference implementation. The Hugging Face checkpoints are a second channel, used mostly through transformers, and are not added to this count.

Capability

3 medium confidence
3.0

Whisper still covers more languages than almost any open recognizer, but on the English board the current open leaders, Qwen3-ASR among them, transcribe with markedly fewer errors.

Verified 2026-09-27