AI Potluck
Back to Gap Map Model components / Speech & audio

Qwen3-TTS

Alibaba Cloud
open weights / Overall score: 2.8

Alibaba's Qwen3-TTS open text-to-speech models (1.7B and 0.6B) built on a 12 Hz speech tokenizer: three-second voice cloning, preset custom voices and voice design from a text description, in 10 languages.

Openness

3 medium confidence
3.0
weights
open(safetensors on the Hub, ungated)
data
closed(no training corpus named)
code
partial(inference package and fine-tuning scripts)
license
Apache-2.0(OSI)

Qwen3-TTS's weights and code are Apache-2.0, OSI-approved. Alibaba publishes fine-tuning scripts but neither its training data nor the recipe behind the released checkpoints.

Adoption

4 high confidence
4.0

Hugging Face downloads of the two most-used checkpoints. The qwen-tts package is marked as not a separate channel because it loads these repositories.

Capability

2 low confidence
2.0

Broad features, voice design and 3-second cloning in 10 languages among them, but the arena entry that most likely corresponds to it sits near the bottom of blind listening tests.

Verified 2026-09-27