VibeVoice
MicrosoftMicrosoft Research's speech models: VibeVoice-ASR transcribes up to an hour of audio in one pass with speaker labels and timestamps in more than 50 languages, with streaming and CPU variants, and the TTS models generate long multi-speaker conversations or stream speech in real time.
Microsoft removed the text-to-speech code from the repository in September 2025 after misuse; the TTS weights remain downloadable, and the entry leads with the ASR line.
Openness
3 medium confidence- weights
- open(ASR and TTS checkpoints on the Hub as safetensors, ungated)
- data
- described(the cards give languages and properties of the training audio only)
- code
- partial(ASR inference and LoRA fine-tuning
- license
- MIT(every checkpoint and the code
Every VibeVoice checkpoint, including the current streaming ASR models, is MIT-licensed and downloads freely. Microsoft describes the training audio without releasing it, and publishes ASR fine-tuning code but not the training pipeline.
- https://huggingface.co/api/models?author=microsoft&search=VibeVoice&limit=100 recorded 2026-09-27
Family listing: eight microsoft/VibeVoice* repositories, all tagged "license:mit", the newest VibeVoice-ASR-Streaming-7B and -1.5B created 2026-09-02.
- https://huggingface.co/api/models/microsoft/VibeVoice-ASR recorded 2026-09-27
Hub record for VibeVoice-ASR: "gated":false and sharded safetensors in the file list.
- https://huggingface.co/microsoft/VibeVoice-1.5B/raw/main/README.md recorded 2026-09-27
Card: "the model is trained only on English and Chinese data"; "The VibeVoice model is limited to research purpose use exploring highly realistic audio dialogue generation".
- https://raw.githubusercontent.com/microsoft/VibeVoice/HEAD/finetuning-asr/README.md recorded 2026-09-27
Fine-tuning README with a "## Training" section using LoRA ("pip install peft").
- https://raw.githubusercontent.com/microsoft/VibeVoice/HEAD/LICENSE recorded 2026-09-27
LICENSE body: "MIT License" / "Copyright (c) 2025 Microsoft".
- https://raw.githubusercontent.com/microsoft/VibeVoice/HEAD/README.md recorded 2026-09-27
README: "we have removed the VibeVoice-TTS code from this repository."; "The VibeVoice-ASR [finetuning code](finetuning-asr/README.md) is now available!"
Adoption
4 high confidenceHugging Face downloads summed over the ASR and TTS checkpoints, split roughly evenly between VibeVoice-ASR and the 1.5B TTS model.
- https://huggingface.co/api/models?author=microsoft&search=VibeVoice&limit=100 recorded 2026-09-27
Trailing-30-day downloads: VibeVoice-ASR 739,222, VibeVoice-1.5B 714,527, VibeVoice-Realtime-0.5B 199,522, VibeVoice-ASR-HF 57,059, VibeVoice-ASR-BitNet 52,719, VibeVoice-ASR-Streaming-1.5B 11,760, VibeVoice-ASR-Streaming-7B 6,732.
Capability
3 medium confidenceMid-table on short English clips, level with Whisper, though its draw is hour-long transcription with speaker labels in a single pass, which that board does not test.
- https://artificialanalysis.ai/text-to-speech/leaderboard recorded 2026-09-27
Leaderboard entry "name":"VibeVoice 1.5B" (creator Microsoft) with "elo":953.25.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
Results CSV row "microsoft/VibeVoice-ASR-HF,5.575".
- https://raw.githubusercontent.com/microsoft/VibeVoice/HEAD/README.md recorded 2026-09-27
README: "a unified streaming ASR model"; "60-minute long-form audio in a single pass".
Verified 2026-09-27