OmniVoice
k2-fsa (Next-gen Kaldi)A massively multilingual zero-shot text-to-speech model from the k2-fsa (Next-gen Kaldi) project: it clones a voice or designs one from a description in more than 600 languages, at up to 40 times faster than real time.
The model's license is stated only in the card's prose, and without a version number.
Openness
2 medium confidence- weights
- open(0.8B safetensors checkpoint with its audio tokenizer on the Hub, ungated)
- data
- partial(581K hours curated from public datasets the paper names
- code
- open(full pipeline in examples/: data preparation, training from scratch, fine-tuning and evaluation)
- license
- CC-BY-NC(the pre-trained model, per the card body (the Hub license field is empty)
The code and the full training pipeline are open, but the card licenses the model itself under CC-BY-NC because of the terms of its training data, so it may not be used commercially. The training set is built from public datasets the paper names rather than published as one.
- https://arxiv.org/abs/2604.00688 recorded 2026-09-27
Abstract: "By leveraging a 581k-hour multilingual dataset curated entirely from open-source data".
- https://huggingface.co/api/models/k2-fsa/OmniVoice recorded 2026-09-27
Hub record: "gated":false, model.safetensors and audio_tokenizer/model.safetensors in the file list, and no license tag.
- https://huggingface.co/k2-fsa/OmniVoice/raw/main/audio_tokenizer/LICENSE recorded 2026-09-27
Bundled tokenizer license: "BOSON HIGGS AUDIO 2 COMMUNITY LICENSE AGREEMENT".
- https://huggingface.co/k2-fsa/OmniVoice/raw/main/README.md recorded 2026-09-27
Card: "Our code is released under the Apache 2.0 License. The pre-trained model is licensed under the CC-BY-NC due to constraints from its training data (e.g., Emilia)."
- https://raw.githubusercontent.com/k2-fsa/OmniVoice/HEAD/examples/README.md recorded 2026-09-27
Examples README: "| Training from scratch | [run_emilia.sh](run_emilia.sh) | Full pipeline on the Emilia dataset (data check, tokenization, training) |".
- https://raw.githubusercontent.com/k2-fsa/OmniVoice/HEAD/LICENSE recorded 2026-09-27
LICENSE body: "Apache License" / "Version 2.0, January 2004".
Adoption
4 high confidenceHugging Face downloads of the model repository. The omnivoice package is marked as not a separate channel because it loads these weights from the Hub.
- https://huggingface.co/api/models?author=k2-fsa&search=OmniVoice&limit=100 recorded 2026-09-27
Trailing-30-day downloads: k2-fsa/OmniVoice 1,378,254.
- https://raw.githubusercontent.com/k2-fsa/OmniVoice/HEAD/README.md recorded 2026-09-27
README usage: "from_pretrained("k2-fsa/OmniVoice"".
Capability
3 low confidenceThe broadest language coverage of any open voice-cloning model. It has no arena rating, so it is placed on its own results on the shared Seed-TTS test sets, level with Kokoro.
- https://arxiv.org/html/2604.00688 recorded 2026-09-27
Paper: "We evaluate OmniVoice on four benchmarks"; "A bilingual (Chinese/English) zero-shot benchmark." for Seed-TTS
- https://huggingface.co/k2-fsa/OmniVoice/raw/main/README.md recorded 2026-09-27
Card: "OmniVoice is a massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages."
Verified 2026-09-27