AI Potluck
Back to Gap Map Model components / Speech & audio

OpenAI speech models

OpenAI
closed / Overall score: n/a

OpenAI's hosted speech models: gpt-transcribe and the gpt-4o transcription models (including a diarizing variant) for recognition, gpt-4o-mini-tts for instructable text-to-speech, and whisper-1 for timestamps and translation, all through the Audio API.

A closed comparator. The open Whisper checkpoints are a separate product, whisper.

Openness

1 high confidence
1.0
weights
closed
data
closed
code
closed
license
Proprietary(API only)

OpenAI's speech models are API-only; no weights, training data or code are published for them.

Adoption

not assessed

Nothing to count: OpenAI publishes no request or developer count for its audio endpoints, and the models have no public download channel.

Capability

4 medium confidence
4.0

Accurate and broad recognition with streaming and diarization, a step behind the most accurate commercial recognizers.

Verified 2026-09-27