Voxtral
Mistral AIMistral AI's speech models: Voxtral Mini 4B Realtime is a natively streaming recognizer for 13 languages with a configurable transcription delay of up to 2.4 s, Voxtral Mini 3B transcribes and answers questions about audio, and Voxtral TTS speaks nine languages with voice adaptation for voice agents.
The 24B Voxtral Small chat model is not part of this entry; it is a chat-first audio language model.
Openness
3 medium confidence- weights
- open(Voxtral Mini 4B Realtime and Voxtral Mini 3B as safetensors on the Hub, ungated)
- data
- closed(the cards describe no training data)
- code
- partial(inference through vLLM, Transformers and mistral-common
- license
- Apache-2.0(the speech-recognition line, Voxtral Mini 4B Realtime 2602 and Voxtral Mini 3B 2507)
- separate-line
- Voxtral 4B TTS 2603, CC-BY-NC-4.0(the text-to-speech model, announced separately with its own name, card and version, and licensed CC-BY-NC-4.0 because it inherits the license of its bundled reference voices
- api-only-tier
- Voxtral Mini Transcribe 2.0(the offline transcription model the Realtime card compares against
Mistral's current speech-recognition models, Voxtral Mini 4B Realtime and Mini 3B, download freely under Apache 2.0. Voxtral 4B TTS is a separate line, announced on its own with its own card, and licensed CC-BY-NC-4.0 because of the reference voices it ships with, so it may not be used commercially; its terms do not carry over to the recognition models. The cards describe no training data and only inference code is published.
- https://huggingface.co/api/models?author=mistralai&search=Voxtral&limit=100&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Family listing: Voxtral-Mini-4B-Realtime-2602, Voxtral-Mini-3B-2507 and Voxtral-Small-24B-2507 tagged "license:apache-2.0"; Voxtral-4B-TTS-2603 tagged "license:cc-by-nc-4.0"; all ungated.
- https://huggingface.co/api/models/mistralai/Voxtral-4B-TTS-2603 recorded 2026-09-27
Hub record: "gated":false, with consolidated.safetensors and voice_embedding files in the file list.
- https://huggingface.co/api/models/mistralai/Voxtral-Mini-3B-2507 recorded 2026-09-27
Hub record: "gated":false, with consolidated.safetensors and sharded safetensors in the file list.
- https://huggingface.co/api/models/mistralai/Voxtral-Mini-4B-Realtime-2602 recorded 2026-09-27
Hub record: "gated":false, with model.safetensors in the file list.
- https://huggingface.co/mistralai/Voxtral-4B-TTS-2603/raw/main/README.md recorded 2026-09-27
Card: "These voices are licensed under CC BY-NC 4, which is the license that the model inherits."; "Thus, this model inherits the same license."
- https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602/raw/main/README.md recorded 2026-09-27
Card: "This model is released in **BF16** under the **Apache-2 license**, ensuring flexibility for both research and commercial use."; the comparison table lists "Voxtral Mini Transcribe 2.0" as an offline model.
- https://mistral.ai/news/voxtral-transcribe-2 recorded 2026-09-28
"Voxtral transcribes at the speed of sound", the recognition line's announcement: "The family includes Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live applications". It does not mention the TTS model.
- https://mistral.ai/news/voxtral-tts recorded 2026-09-28
"Speaking of Voxtral", Mistral's own announcement of the TTS model: "A model with several reference voices is available as open weights on Hugging Face under CC BY NC 4.0 license", and "It works alongside Voxtral Transcribe for full speech-to-speech, or integrates into any existing speech-to-text and LLM stack".
Adoption
4 high confidenceHugging Face downloads summed over the three declared checkpoints, almost all of them the streaming recognizer.
- https://huggingface.co/api/models?author=mistralai&search=Voxtral&limit=100&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: Voxtral-Mini-4B-Realtime-2602 1,765,082, Voxtral-Mini-3B-2507 247,167, Voxtral-4B-TTS-2603 1,583.
Capability
3 medium confidenceMid-table on both boards, level with Whisper for recognition and close to Kokoro for synthesis; the realtime model trades some accuracy for streaming with a sub-second delay.
- https://artificialanalysis.ai/text-to-speech/leaderboard recorded 2026-09-27
Leaderboard entry "name":"Voxtral TTS" (creator Mistral) with "openWeights":true, ranked 42nd of the 91 models on the board.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV rows "mistralai/Voxtral-Mini-3B-2507,5.53875" and "mistralai/Voxtral-Mini-4B-Realtime-2602,6.4625".
- https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602/raw/main/README.md recorded 2026-09-27
Card: "Built with a **natively streaming architecture** and a custom causal audio encoder - it allows configurable transcription delays (240ms to 2.4s)".
Verified 2026-09-27