AI Potluck
Back to Gap Map Model components / Speech & audio

XTTS

Coqui
restricted / Overall score: 2.8

Coqui's voice-cloning text-to-speech model: XTTS-v2 clones a voice from a six-second sample and speaks it in 17 languages, including across languages, with streaming output. It runs on the Coqui TTS engine, now maintained by Idiap as the coqui-tts package.

Coqui, the company, has shut down: its website no longer resolves and the original repository is unmaintained, so XTTS-v2 is its last release.

Openness

2 high confidence
2.0
weights
open(XTTS-v2 .pth checkpoints on the Hub, ungated)
data
closed(no training-data statement on either card)
code
partial(inference plus a GPT-encoder fine-tuning recipe in the maintained Idiap fork
license
Coqui-Public-Model-License-1.0.0(XTTS-v2 and XTTS-v1 weights

XTTS-v2 downloads freely, but the Coqui Public Model License allows only non-commercial use of the model and its outputs. Nothing is published about the training data, and only a fine-tuning recipe for one component is available. That license permits no commercial use, so its tier is commercial_forbidden.

Adoption

4 high confidence
4.0

Hugging Face downloads of the XTTS checkpoints; installs through the Coqui TTS packages fetch the same files.

Capability

2 medium confidence
2.0

Still the most downloaded open voice-cloning model, but listeners in blind tests rank it near the bottom of the board, a step below Kokoro.

Verified 2026-09-27