AI Potluck
Back to Gap Map Model components / Speech & audio

F5-TTS

SWivid
restricted / Overall score: 3.0

A flow-matching diffusion-transformer text-to-speech model for zero-shot voice cloning in English and Chinese, released with its training and fine-tuning code. The project is led by SJTU X-LANCE researchers under the SWivid account.

The code is MIT; the pretrained weights are CC-BY-NC-4.0 because they were trained on the Emilia dataset.

Openness

2 high confidence
2.0
weights
open(checkpoints on the Hub, ungated)
data
open(trained on the public Emilia dataset)
code
open(training and fine-tuning code in the repository)
license
CC-BY-NC-4.0(weights

F5-TTS publishes its code under MIT and trains on the public Emilia corpus, but the weights carry a non-commercial license inherited from that data. Downloadable is not usable commercially, and that is what sets the score.

Adoption

3 high confidence
3.0

Hugging Face downloads of the weights repository. The f5-tts package is marked as not a separate channel because it loads this repository.

Capability

3 low confidence
3.0

An influential open voice-cloning architecture, limited to two languages. It is not on the Artificial Analysis TTS leaderboard, so it is placed on its own reported results on a shared public test set, level with Kokoro, rather than on a blind listening test.

  • https://arxiv.org/pdf/2410.06885 recorded 2026-09-27

    Paper (PDF; read with a text extractor): Table 6 gives evaluation results of F5-TTS on LibriSpeech-PC test-clean, Seed-TTS test-en and Seed-TTS test-zh, and Table 2 compares it with other systems on the two Seed-TTS test sets.

Verified 2026-09-27