CosyVoice
Alibaba CloudAlibaba's multilingual voice generation line (CosyVoice, CosyVoice 2 and Fun-CosyVoice 3), with zero-shot voice cloning and instruction control over dialect, emotion and speed in 9 languages and more than 18 Chinese dialects. The repository ships inference, training and deployment code.
The repository moved from FunAudioLLM/CosyVoice to QwenAudio/CosyVoice; both names resolve to the same repository. Code and weights are Apache-2.0.
Openness
3 medium confidence- weights
- open(checkpoints on the Hub, ungated)
- data
- closed(no training corpus named in the card or README)
- code
- open(inference, training and deployment code in the repository)
- license
- Apache-2.0(OSI)
CosyVoice's code and weights are Apache-2.0, OSI-approved, and the repository includes training code. Alibaba does not publish or name the data the checkpoints were trained on, which keeps it at open weights.
- https://huggingface.co/api/models/FunAudioLLM/Fun-CosyVoice3-0.5B-2512 recorded 2026-09-27
"license:apache-2.0" tag and "gated":false.
- https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512/raw/main/README.md recorded 2026-09-27
Card front matter "license: apache-2.0"; the card reports benchmark results and usage but names no training dataset.
- https://raw.githubusercontent.com/QwenAudio/CosyVoice/HEAD/LICENSE recorded 2026-09-27
LICENSE body: "Apache License" "Version 2.0, January 2004".
- https://ungh.cc/repos/QwenAudio/CosyVoice recorded 2026-09-27
Repository description: "Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability."
Adoption
3 high confidenceHugging Face downloads of the two current checkpoints; the older CosyVoice 1 repositories add little.
- https://huggingface.co/api/models?author=FunAudioLLM&search=CosyVoice recorded 2026-09-27
Fun-CosyVoice3-0.5B-2512 193,282 and CosyVoice2-0.5B 4,510 downloads in the trailing 30 days.
Capability
3 low confidenceA leading open Chinese voice-cloning line with dialect coverage few peers match. It is not on the Artificial Analysis TTS leaderboard, so it is placed on its own reported results on a shared public test set, level with Kokoro, rather than on a blind listening test.
- https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512/raw/main/README.md recorded 2026-09-27
Results table columns "test-zh<br>CER (%) ↓" and "test-en<br>WER (%) ↓" on the Seed-TTS evaluation.
- https://raw.githubusercontent.com/QwenAudio/CosyVoice/HEAD/README.md recorded 2026-09-27
"Covers 9 common languages (Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian), 18+ Chinese dialects/accents".
Verified 2026-09-27