AI Potluck
Back to Gap Map Model components / Speech & audio

Zonos

Zyphra
open weights / Overall score: 3.0

Zyphra's expressive text-to-speech models with voice cloning and emotion control: ZONOS2, an 8B mixture-of-experts model trained on more than six million hours of speech, streams in three tiers of languages, and the earlier Zonos v0.1 outputs 44 kHz speech.

The ZONOS2 weights are labeled Apache-2.0 on the Hub and in Zyphra's announcement, while its code repository's LICENSE is MIT.

Openness

3 medium confidence
3.0
weights
open(ZONOS2 (model.pth) and GGUF conversions on the Hub, ungated)
data
described(more than six million hours of web-scale multilingual speech
code
partial(inference server and offline API
license
Apache-2.0(ZONOS2 weights per the Hub card and Zyphra's announcement
superseded-release
Zonos-v0.1(2025 transformer and hybrid models, Apache-2.0, still published)

ZONOS2, the current release, downloads freely under Apache-2.0, as did Zonos v0.1 before it. Zyphra describes its six million hours of training audio without naming or releasing them and publishes only inference code.

Adoption

3 medium confidence
3.0

Hugging Face downloads summed over the declared checkpoints, still mostly the 2025 v0.1 transformer model.

Capability

3 low confidence
3.0

The new ZONOS2 has no listener rating yet, so it is placed on its own results on the shared Seed-TTS test set, level with Kokoro. Its predecessor sits in the lower part of the arena, and a rating for ZONOS2 would replace this reading.

Verified 2026-09-27