AI Potluck
Back to Gap Map Model components / Speech & audio

OmniVoice

k2-fsa (Next-gen Kaldi)
restricted / Overall score: 3.4

A massively multilingual zero-shot text-to-speech model from the k2-fsa (Next-gen Kaldi) project: it clones a voice or designs one from a description in more than 600 languages, at up to 40 times faster than real time.

The model's license is stated only in the card's prose, and without a version number.

Openness

2 medium confidence
2.0
weights
open(0.8B safetensors checkpoint with its audio tokenizer on the Hub, ungated)
data
partial(581K hours curated from public datasets the paper names
code
open(full pipeline in examples/: data preparation, training from scratch, fine-tuning and evaluation)
license
CC-BY-NC(the pre-trained model, per the card body (the Hub license field is empty)

The code and the full training pipeline are open, but the card licenses the model itself under CC-BY-NC because of the terms of its training data, so it may not be used commercially. The training set is built from public datasets the paper names rather than published as one.

Adoption

4 high confidence
4.0

Hugging Face downloads of the model repository. The omnivoice package is marked as not a separate channel because it loads these weights from the Hub.

Capability

3 low confidence
3.0

The broadest language coverage of any open voice-cloning model. It has no arena rating, so it is placed on its own results on the shared Seed-TTS test sets, level with Kokoro.

Verified 2026-09-27