AI Potluck
Back to Gap Map Model components / Speech & audio

Moshi

Kyutai
open weights / Overall score: 4.2(strong)

Kyutai's speech-text foundation model and full-duplex spoken dialogue system: it listens and speaks at the same time, with about 200 ms of practical latency, using the Mimi streaming audio codec. Released as the Moshiko and Moshika voices for PyTorch, MLX and Rust.

Weights are CC-BY-4.0; the Python code is MIT and the Rust backend Apache-2.0. One gated research checkpoint, moshika-rl-seamless, is CC-BY-NC-4.0 and is not in the release the README lists.

Openness

3 high confidence
3.0
weights
open(PyTorch, MLX and Candle checkpoints, ungated)
data
described(7M hours of audio, Fisher and a Kyutai-collected set
code
partial(inference, server and a separate fine-tuning repository)
license
CC-BY-4.0(Moshi v0.1 release weights)
research-checkpoint
moshika-rl-seamless(2026, gated CC-BY-NC-4.0 fine-tune published with the Interactivity Alignment paper

Moshi's current release, the v0.1 Moshiko and Moshika checkpoints for PyTorch, MLX and Candle, is CC-BY-4.0, an attribution-only license, with code under MIT and Apache-2.0. Kyutai describes its 7 million hours of training audio without releasing them and publishes inference code and a separate fine-tuning repository rather than the training pipeline. The gated, non-commercial moshika-rl-seamless checkpoint belongs to Kyutai's separate Interactivity Alignment research release and does not set the score.

Adoption

3 medium confidence
3.0

Hugging Face downloads of the three most-used checkpoints. The moshi package is marked as not a separate channel because it loads these repositories.

Capability

5 medium confidence
5.0

Moshi holds a live two-way spoken conversation, which no recognizer or synthesizer here does, and it was the first open model to do so in real time.

Verified 2026-09-27