Moshi
KyutaiKyutai's speech-text foundation model and full-duplex spoken dialogue system: it listens and speaks at the same time, with about 200 ms of practical latency, using the Mimi streaming audio codec. Released as the Moshiko and Moshika voices for PyTorch, MLX and Rust.
Weights are CC-BY-4.0; the Python code is MIT and the Rust backend Apache-2.0. One gated research checkpoint, moshika-rl-seamless, is CC-BY-NC-4.0 and is not in the release the README lists.
Openness
3 high confidence- weights
- open(PyTorch, MLX and Candle checkpoints, ungated)
- data
- described(7M hours of audio, Fisher and a Kyutai-collected set
- code
- partial(inference, server and a separate fine-tuning repository)
- license
- CC-BY-4.0(Moshi v0.1 release weights)
- research-checkpoint
- moshika-rl-seamless(2026, gated CC-BY-NC-4.0 fine-tune published with the Interactivity Alignment paper
Moshi's current release, the v0.1 Moshiko and Moshika checkpoints for PyTorch, MLX and Candle, is CC-BY-4.0, an attribution-only license, with code under MIT and Apache-2.0. Kyutai describes its 7 million hours of training audio without releasing them and publishes inference code and a separate fine-tuning repository rather than the training pipeline. The gated, non-commercial moshika-rl-seamless checkpoint belongs to Kyutai's separate Interactivity Alignment research release and does not set the score.
- https://huggingface.co/api/collections/kyutai/interactivity-alignment-6a1fe5f14879ce66956362e2 recorded 2026-09-27
Collection "Interactivity Alignment", "Full-duplex speech models post-trained with reinforcement learning for improved conversational interactivity.": paper 2606.11167, kyutai/moshika-rl-seamless, kyutai/personaplex-rl-seamless and a samples dataset.
- https://huggingface.co/api/collections/kyutai/moshi-v01-release-66eaeaf3302bef6bd9ad7acd recorded 2026-09-27
Collection "Moshi v0.1 Release", "MLX, Candle & PyTorch model checkpoints released as part of the Moshi release from Kyutai": paper 2410.00037, kyutai/mimi and the moshiko and moshika pytorch, mlx and candle checkpoints; moshika-rl-seamless is not among them.
- https://huggingface.co/api/models/kyutai/moshika-rl-seamless recorded 2026-09-27
kyutai/moshika-rl-seamless: "license:cc-by-nc-4.0", "gated":"auto", "base_model:finetune:kyutai/moshika-pytorch-bf16", "arxiv:2606.11167", "createdAt":"2026-06-02T08:37:38.000Z".
- https://huggingface.co/api/models/kyutai/moshiko-pytorch-bf16 recorded 2026-09-27
"license:cc-by-4.0" tag, "gated":false, with model.safetensors in the file list.
- https://huggingface.co/kyutai/moshiko-pytorch-bf16/raw/main/README.md recorded 2026-09-27
Training Data section: "Unsupervised audio dataset:** used for pre-training, this is a collection of 7 million hours of readily available audio content"; "A dataset of 170 hours of natural and scripted conversation between multiple pairs of participants, collected by Kyutai."
- https://raw.githubusercontent.com/kyutai-labs/moshi/HEAD/README.md recorded 2026-09-27
"The weights for the models are released under the CC-BY 4.0 license."; "If you want to fine tune Moshi, head out to [kyutai-labs/moshi-finetune]".
Adoption
3 medium confidenceHugging Face downloads of the three most-used checkpoints. The moshi package is marked as not a separate channel because it loads these repositories.
- https://huggingface.co/api/models?author=kyutai&search=moshi recorded 2026-09-27
Trailing-30-day downloads: moshiko-pytorch-bf16 179,111, moshiko-candle-q8 22,422, moshika-pytorch-bf16 4,342.
Capability
5 medium confidenceMoshi holds a live two-way spoken conversation, which no recognizer or synthesizer here does, and it was the first open model to do so in real time.
- https://huggingface.co/kyutai/moshiko-pytorch-bf16/raw/main/README.md recorded 2026-09-27
"Moshi is the first real-time full-duplex spoken large language model".
- https://raw.githubusercontent.com/kyutai-labs/moshi/HEAD/README.md recorded 2026-09-27
"is a speech-text foundation model and **full-duplex** spoken dialogue framework"; "with a practical overall latency as low as 200ms on an L4 GPU."
Verified 2026-09-27