Whisper
OpenAIOpenAI's open-weight speech recognition and translation model family, from tiny to large-v3 and the faster large-v3-turbo. It transcribes about 99 languages and translates speech into English.
The large-v3 checkpoints carry Apache-2.0 on the Hub and large-v3-turbo carries MIT, the same license as the openai/whisper code. OpenAI's hosted transcription models are a separate product, openai-speech.
Openness
3 high confidence- weights
- open(safetensors checkpoints on the Hub, ungated)
- data
- described(1M hours weakly labeled plus 4M hours pseudo-labeled audio, not released)
- code
- partial(inference and decoding code only)
- license
- MIT(large-v3-turbo weights and the code)+Apache-2.0(large-v3 and earlier weights)
Whisper's weights download without a gate under MIT and Apache-2.0, both OSI-approved. OpenAI describes the five million hours of audio it trained on but has not released them, and the repository holds inference code rather than a training pipeline, so this is open weights rather than a reproducible model.
- https://huggingface.co/api/models/openai/whisper-large-v3-turbo recorded 2026-09-27
"license:mit" tag, "gated":false, and model.safetensors in the file list.
- https://huggingface.co/openai/whisper-large-v3/raw/main/README.md recorded 2026-09-27
Card front matter "license: apache-2.0"; "The large-v3 checkpoint is trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio collected using Whisper large-v2." No dataset is linked.
- https://raw.githubusercontent.com/openai/whisper/HEAD/LICENSE recorded 2026-09-27
Repository LICENSE body: "MIT License" / "Copyright (c) 2022 OpenAI".
- https://raw.githubusercontent.com/openai/whisper/HEAD/README.md recorded 2026-09-27
The README covers installation, command-line and Python transcription with the released checkpoints; it documents no training procedure.
Adoption
4 high confidenceRead from PyPI, the route that leads for a product declaring a package: installs of openai-whisper, the reference implementation. The Hugging Face checkpoints are a second channel, used mostly through transformers, and are not added to this count.
- https://huggingface.co/api/models?author=openai&search=whisper recorded 2026-09-27
Twelve openai/whisper-* repositories; downloads include openai/whisper-large-v3-turbo 6,477,272 and openai/whisper-large-v3 4,600,554 in the trailing 30 days.
- https://pypistats.org/api/packages/openai-whisper/recent recorded 2026-09-27
"last_month":5190666 for openai-whisper.
Capability
3 medium confidenceWhisper still covers more languages than almost any open recognizer, but on the English board the current open leaders, Qwen3-ASR among them, transcribe with markedly fewer errors.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV rows "openai/whisper-large-v3,5.78" and "openai/whisper-large-v3-turbo,6.3575"; both list 99 languages.
Verified 2026-09-27