AI Potluck
Back to Gap Map Model components / Speech & audio

SeamlessM4T

Meta
restricted / Overall score: 2.4

Meta's multilingual speech translation model: one model translates speech to speech, speech to text and text to speech, and transcribes, with speech input in about 100 languages and speech output in 35. The v2 release adds a streaming variant for simultaneous translation.

Openness

2 high confidence
2.0
weights
open(SeamlessM4T v2 Large as fairseq2 and Transformers checkpoints, ungated)
data
partial(SeamlessAlign metadata is published and rebuildable
code
partial(inference, evaluation and a demonstration fine-tuning script)
license
CC-BY-NC-4.0(SeamlessM4T v1 and v2 and SeamlessStreaming weights
sibling-model
seamless-expressive(SeamlessExpressive, released with v2 under the Seamless Licensing Agreement (noncommercial research only) behind a request form

SeamlessM4T v2 downloads freely, but its weights are CC-BY-NC-4.0 and may not be used commercially; only the code and the speech encoder are MIT. Meta publishes the metadata to rebuild its aligned training corpus, not the full training mixture.

Adoption

3 medium confidence
3.0

Hugging Face downloads of the Transformers-format checkpoints; the fairseq2 repositories report none because the Hub does not count their files.

Capability

2 low confidence
2.0

The only open model on the map that translates speech directly into speech in dozens of languages. Its results sit in downloadable metrics files rather than on a public board, so it sits a step below Whisper.

Verified 2026-09-27