AI Potluck
Back to Gap Map Model components / Speech & audio

VibeVoice

Microsoft
open weights / Overall score: 3.4

Microsoft Research's speech models: VibeVoice-ASR transcribes up to an hour of audio in one pass with speaker labels and timestamps in more than 50 languages, with streaming and CPU variants, and the TTS models generate long multi-speaker conversations or stream speech in real time.

Microsoft removed the text-to-speech code from the repository in September 2025 after misuse; the TTS weights remain downloadable, and the entry leads with the ASR line.

Openness

3 medium confidence
3.0
weights
open(ASR and TTS checkpoints on the Hub as safetensors, ungated)
data
described(the cards give languages and properties of the training audio only)
code
partial(ASR inference and LoRA fine-tuning
license
MIT(every checkpoint and the code

Every VibeVoice checkpoint, including the current streaming ASR models, is MIT-licensed and downloads freely. Microsoft describes the training audio without releasing it, and publishes ASR fine-tuning code but not the training pipeline.

Adoption

4 high confidence
4.0

Hugging Face downloads summed over the ASR and TTS checkpoints, split roughly evenly between VibeVoice-ASR and the 1.5B TTS model.

Capability

3 medium confidence
3.0

Mid-table on short English clips, level with Whisper, though its draw is hour-long transcription with speaker labels in a single pass, which that board does not test.

Verified 2026-09-27