AI Potluck
Back to Gap Map Model components / Speech & audio

CosyVoice

Alibaba Cloud
open weights / Overall score: 3.0

Alibaba's multilingual voice generation line (CosyVoice, CosyVoice 2 and Fun-CosyVoice 3), with zero-shot voice cloning and instruction control over dialect, emotion and speed in 9 languages and more than 18 Chinese dialects. The repository ships inference, training and deployment code.

The repository moved from FunAudioLLM/CosyVoice to QwenAudio/CosyVoice; both names resolve to the same repository. Code and weights are Apache-2.0.

Openness

3 medium confidence
3.0
weights
open(checkpoints on the Hub, ungated)
data
closed(no training corpus named in the card or README)
code
open(inference, training and deployment code in the repository)
license
Apache-2.0(OSI)

CosyVoice's code and weights are Apache-2.0, OSI-approved, and the repository includes training code. Alibaba does not publish or name the data the checkpoints were trained on, which keeps it at open weights.

Adoption

3 high confidence
3.0

Hugging Face downloads of the two current checkpoints; the older CosyVoice 1 repositories add little.

Capability

3 low confidence
3.0

A leading open Chinese voice-cloning line with dialect coverage few peers match. It is not on the Artificial Analysis TTS leaderboard, so it is placed on its own reported results on a shared public test set, level with Kokoro, rather than on a blind listening test.

Verified 2026-09-27