SeamlessM4T
MetaMeta's multilingual speech translation model: one model translates speech to speech, speech to text and text to speech, and transcribes, with speech input in about 100 languages and speech output in 35. The v2 release adds a streaming variant for simultaneous translation.
Openness
2 high confidence- weights
- open(SeamlessM4T v2 Large as fairseq2 and Transformers checkpoints, ungated)
- data
- partial(SeamlessAlign metadata is published and rebuildable
- code
- partial(inference, evaluation and a demonstration fine-tuning script)
- license
- CC-BY-NC-4.0(SeamlessM4T v1 and v2 and SeamlessStreaming weights
- sibling-model
- seamless-expressive(SeamlessExpressive, released with v2 under the Seamless Licensing Agreement (noncommercial research only) behind a request form
SeamlessM4T v2 downloads freely, but its weights are CC-BY-NC-4.0 and may not be used commercially; only the code and the speech encoder are MIT. Meta publishes the metadata to rebuild its aligned training corpus, not the full training mixture.
- https://huggingface.co/api/models/facebook/seamless-m4t-v2-large?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=tags&expand[]=siblings recorded 2026-09-27
Hub record for seamless-m4t-v2-large: "gated":false, tag "license:cc-by-nc-4.0", and seamlessM4T_v2_large.pt with safetensors shards.
- https://raw.githubusercontent.com/facebookresearch/seamless_communication/main/LICENSE recorded 2026-09-27
LICENSE body: "Attribution-NonCommercial 4.0 International".
- https://raw.githubusercontent.com/facebookresearch/seamless_communication/main/README.md recorded 2026-09-27
README: "We have three license categories."; "The following models are CC-BY-NC 4.0 licensed as found in the [LICENSE](LICENSE):" covering "SeamlessM4T models (v1 and v2)."; "We open-source the metadata to SeamlessAlign".
- https://raw.githubusercontent.com/facebookresearch/seamless_communication/main/SEAMLESS_LICENSE recorded 2026-09-27
SEAMLESS_LICENSE, titled "Seamless Licensing Agreement", grants use "solely for Noncommercial Research Uses".
- https://raw.githubusercontent.com/facebookresearch/seamless_communication/main/src/seamless_communication/cli/m4t/finetune/README.md recorded 2026-09-27
Fine-tuning README: "The trainer and dataloader were designed mainly for demonstration purposes."
Adoption
3 medium confidenceHugging Face downloads of the Transformers-format checkpoints; the fairseq2 repositories report none because the Hub does not count their files.
- https://huggingface.co/api/models?author=facebook&search=seamless&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: seamless-m4t-v2-large 299,813, hf-seamless-m4t-medium 114,820, hf-seamless-m4t-large 718, seamless-streaming 0.
Capability
2 low confidenceThe only open model on the map that translates speech directly into speech in dozens of languages. Its results sit in downloadable metrics files rather than on a public board, so it sits a step below Whisper.
- https://artificialanalysis.ai/speech-to-text recorded 2026-09-27
The speech-to-text table lists no row for this model.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
The results CSV lists no facebook/seamless row among its 71 models.
- https://huggingface.co/facebook/seamless-m4t-v2-large/raw/main/README.md recorded 2026-09-27
Card: "Speech-to-speech translation (S2ST)"; "101 languages for speech input."; "35 languages for speech output."; "The evaluation data ids for FLEURS, CoVoST2 and CVSS-C can be found".
Verified 2026-09-27