AI Potluck
Back to Gap Map Model components / Speech & audio

Kokoro

hexgrad
open weights / Overall score: 3.8

An 82-million-parameter open-weight text-to-speech model with dozens of preset voices across English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin. Small enough to run on a CPU, it is widely deployed in local apps and commercial APIs.

Maintained by the individual developer hexgrad; the kokoro Python package is the inference library.

Openness

3 high confidence
3.0
weights
open(.pth checkpoint and voice packs, ungated)
data
described(a few hundred hours of permissively licensed audio, listed in part)
code
partial(inference library only)
license
Apache-2.0(OSI)

Kokoro's weights and inference library are Apache-2.0, an OSI-approved license. The card lists some of the openly licensed audio it trained on and sizes the rest at a few hundred hours, but the training set and training code are not published.

Adoption

5 high confidence
5.0

Hugging Face downloads of the single weights repository. The kokoro package is marked as not a separate channel because it loads this repository.

Capability

3 medium confidence
3.0

Kokoro is mid-table in blind listening tests, well behind the commercial leaders, but it is by far the smallest model near that position.

Verified 2026-09-27