AI Potluck
Back to Gap Map Model components / Speech & audio

Voxtral

Mistral AI
open weights / Overall score: 3.4

Mistral AI's speech models: Voxtral Mini 4B Realtime is a natively streaming recognizer for 13 languages with a configurable transcription delay of up to 2.4 s, Voxtral Mini 3B transcribes and answers questions about audio, and Voxtral TTS speaks nine languages with voice adaptation for voice agents.

The 24B Voxtral Small chat model is not part of this entry; it is a chat-first audio language model.

Openness

3 medium confidence
3.0
weights
open(Voxtral Mini 4B Realtime and Voxtral Mini 3B as safetensors on the Hub, ungated)
data
closed(the cards describe no training data)
code
partial(inference through vLLM, Transformers and mistral-common
license
Apache-2.0(the speech-recognition line, Voxtral Mini 4B Realtime 2602 and Voxtral Mini 3B 2507)
separate-line
Voxtral 4B TTS 2603, CC-BY-NC-4.0(the text-to-speech model, announced separately with its own name, card and version, and licensed CC-BY-NC-4.0 because it inherits the license of its bundled reference voices
api-only-tier
Voxtral Mini Transcribe 2.0(the offline transcription model the Realtime card compares against

Mistral's current speech-recognition models, Voxtral Mini 4B Realtime and Mini 3B, download freely under Apache 2.0. Voxtral 4B TTS is a separate line, announced on its own with its own card, and licensed CC-BY-NC-4.0 because of the reference voices it ships with, so it may not be used commercially; its terms do not carry over to the recognition models. The cards describe no training data and only inference code is published.

Adoption

4 high confidence
4.0

Hugging Face downloads summed over the three declared checkpoints, almost all of them the streaming recognizer.

Capability

3 medium confidence
3.0

Mid-table on both boards, level with Whisper for recognition and close to Kokoro for synthesis; the realtime model trades some accuracy for streaming with a sub-second delay.

Verified 2026-09-27