AI Potluck
Back to Gap Map Model components / Speech & audio

CSM (Conversational Speech Model)

Sesame
open weights / Overall score: 1.8

Sesame's Conversational Speech Model: a speech generator with a Llama backbone and a small audio decoder producing Mimi codes, which conditions on the preceding conversation, text and audio, to speak with context-appropriate prosody. The released CSM-1B is a base model without a fixed voice.

The repository has not been pushed to since May 2025; only the smallest of the three sizes Sesame trained is released.

Openness

3 medium confidence
3.0
weights
open(csm-1b safetensors behind an automatically approved contact-information click-through)
data
described(about one million hours of publicly available, mostly English audio, transcribed and filtered
code
partial(inference code
license
Apache-2.0(code and weights
unreleased-sizes
CSM Small and Medium(the 3B and 8B models the research post evaluates were never distributed

CSM-1B is Apache-2.0 and downloads after a contact-details click-through. Sesame describes its million hours of training audio without releasing it, and publishes inference code only. The larger sizes it evaluated were never released and do not set the score.

Adoption

3 high confidence
3.0

Hugging Face downloads of the single released checkpoint.

Capability

1 low confidence
1.0

The model behind a widely noticed voice demo, but the released base model has no public rating and the published results describe a larger model that was never released, so it sits well below Kokoro.

Verified 2026-09-27