OpenAI speech models
OpenAIOpenAI's hosted speech models: gpt-transcribe and the gpt-4o transcription models (including a diarizing variant) for recognition, gpt-4o-mini-tts for instructable text-to-speech, and whisper-1 for timestamps and translation, all through the Audio API.
A closed comparator. The open Whisper checkpoints are a separate product, whisper.
Openness
1 high confidence- weights
- closed
- data
- closed
- code
- closed
- license
- Proprietary(API only)
OpenAI's speech models are API-only; no weights, training data or code are published for them.
- https://developers.openai.com/api/docs/guides/speech-to-text.md recorded 2026-09-27
Guide documents the hosted "/v1/audio/transcriptions" endpoint with "gpt-transcribe"; models are called over the API with a key and none is downloadable.
- https://developers.openai.com/api/docs/guides/text-to-speech.md recorded 2026-09-27
Text-to-speech is the hosted "gpt-4o-mini-tts" model called through the API.
Adoption
not assessedNothing to count: OpenAI publishes no request or developer count for its audio endpoints, and the models have no public download channel.
- https://developers.openai.com/api/docs/guides/speech-to-text.md recorded 2026-09-27
Guide documents transcription with "gpt-transcribe", diarization with "gpt-4o-transcribe-diarize" and translation with "whisper-1"; no usage figure appears.
Capability
4 medium confidenceAccurate and broad recognition with streaming and diarization, a step behind the most accurate commercial recognizers.
- https://artificialanalysis.ai/speech-to-text recorded 2026-09-27
Table rows "GPT Transcribe, OpenAI" with "3.3%" and "GPT-4o Transcribe" with "4.0%".
- https://artificialanalysis.ai/text-to-speech/leaderboard recorded 2026-09-27
Entry "name":"TTS-1 HD" with "elo":1103.76, creator OpenAI.
Verified 2026-09-27