CSM (Conversational Speech Model)
SesameSesame's Conversational Speech Model: a speech generator with a Llama backbone and a small audio decoder producing Mimi codes, which conditions on the preceding conversation, text and audio, to speak with context-appropriate prosody. The released CSM-1B is a base model without a fixed voice.
The repository has not been pushed to since May 2025; only the smallest of the three sizes Sesame trained is released.
Openness
3 medium confidence- weights
- open(csm-1b safetensors behind an automatically approved contact-information click-through)
- data
- described(about one million hours of publicly available, mostly English audio, transcribed and filtered
- code
- partial(inference code
- license
- Apache-2.0(code and weights
- unreleased-sizes
- CSM Small and Medium(the 3B and 8B models the research post evaluates were never distributed
CSM-1B is Apache-2.0 and downloads after a contact-details click-through. Sesame describes its million hours of training audio without releasing it, and publishes inference code only. The larger sizes it evaluated were never released and do not set the score.
- https://huggingface.co/api/models/sesame/csm-1b recorded 2026-09-27
Hub record: tag "license:apache-2.0", "gated":"auto", and model.safetensors in the file list.
- https://raw.githubusercontent.com/SesameAILabs/csm/main/LICENSE recorded 2026-09-27
LICENSE body: "Apache License" / "Version 2.0, January 2004".
- https://raw.githubusercontent.com/SesameAILabs/csm/main/README.md recorded 2026-09-27
README: "**2025/03/13** - We are releasing the 1B CSM variant."; "A fine-tuned variant of CSM powers the [interactive voice demo](https://www.sesame.com/voicedemo)".
- https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice recorded 2026-09-27
Research post: "After filtering, the dataset consists of approximately one million hours of predominantly English audio."; "8B backbone, 300M decoder" for Medium.
Adoption
3 high confidenceHugging Face downloads of the single released checkpoint.
- https://huggingface.co/api/models/sesame/csm-1b recorded 2026-09-27
"downloads":138170 for sesame/csm-1b.
Capability
1 low confidenceThe model behind a widely noticed voice demo, but the released base model has no public rating and the published results describe a larger model that was never released, so it sits well below Kokoro.
- https://raw.githubusercontent.com/SesameAILabs/csm/main/README.md recorded 2026-09-27
README: "The model open-sourced here is a base generation model. It is capable of producing a variety of voices, but it has not been fine-tuned on any specific voice."
- https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice recorded 2026-09-27
Research post: "Comparative Mean Opinion Score (CMOS) studies using the Expresso dataset to assess the naturalness and prosodic appropriateness of generated speech for CSM-Medium."
Verified 2026-09-27