MMS (Massively Multilingual Speech)
MetaMeta's Massively Multilingual Speech models: one wav2vec 2.0 recognizer with per-language adapters for over 1,100 languages, language identification across about 4,000, and a separate text-to-speech voice for each of more than 1,100 languages.
Meta's later Omnilingual ASR (2025, Apache-2.0) is a separately named system and is not folded into this entry.
Openness
2 high confidence- weights
- open(ASR, language-ID and per-language VITS TTS checkpoints on the Hub, ungated)
- data
- closed(the MMS-lab and unlabeled corpora are not released)
- code
- partial(inference and adapter fine-tuning
- license
- CC-BY-NC-4.0(every MMS checkpoint, code and weights alike)
Every MMS checkpoint downloads freely but is licensed CC-BY-NC-4.0, which rules out commercial use. The labeled and unlabeled corpora it was trained on are not released.
- https://huggingface.co/api/models?author=facebook&search=mms&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Family listing: the facebook/mms-* ASR, LID and TTS checkpoints tagged "license:cc-by-nc-4.0", created between May and September 2023; only the empty facebook/MMS repository carries another tag.
- https://huggingface.co/api/models/facebook/mms-1b-all?expand[]=downloads&expand[]=cardData&expand[]=gated&expand[]=createdAt&expand[]=lastModified&expand[]=tags&expand[]=siblings recorded 2026-09-27
Hub record for mms-1b-all: "gated":false, tag "license:cc-by-nc-4.0", with per-language adapter safetensors files.
- https://huggingface.co/facebook/mms-1b-all/raw/main/README.md recorded 2026-09-27
Card: "has been fine-tuned from [facebook/mms-1b](https://huggingface.co/facebook/mms-1b) on 1162 languages."
- https://raw.githubusercontent.com/facebookresearch/fairseq/main/examples/mms/README.md recorded 2026-09-27
MMS README: "The MMS code and model weights are released under the CC-BY-NC 4.0 license."; the language-coverage table names "MMS-lab"; fine-tuning goes through "the official 🤗 Transformers examples".
Adoption
3 medium confidenceHugging Face downloads of the six declared checkpoints. More than a thousand per-language TTS voices sit outside that figure, so the family as a whole is used more than it shows.
- https://huggingface.co/api/models?author=facebook&search=mms-tts&sort=downloads&direction=-1&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: mms-tts-eng 133,619.
- https://huggingface.co/api/models?author=facebook&search=mms&limit=1000&expand[]=downloads&expand[]=createdAt&expand[]=tags&expand[]=gated recorded 2026-09-27
Trailing-30-day downloads: mms-1b-all 266,668, mms-lid-1024 180,708, mms-lid-126 78,403, mms-300m 30,345, mms-1b 4,534.
Capability
2 low confidenceUnmatched breadth for low-resource languages, with recognition, identification and synthesis in one family. Its accuracy is not compared on any public board, so it sits a step below Whisper, which is.
- https://artificialanalysis.ai/speech-to-text recorded 2026-09-27
The speech-to-text table lists no row for this model.
- https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-results/resolve/d2c5b384deccdb82834f41aeaffcc618c00efa2f/english_short_latest.csv recorded 2026-09-27
The results CSV lists no facebook/mms row among its 71 models.
- https://raw.githubusercontent.com/facebookresearch/fairseq/main/examples/mms/README.md recorded 2026-09-27
README: "building a single multilingual speech recognition model supporting over 1,100 languages (more than 10 times as many as before), language identification models able to identify over [4,000 languages]".
Verified 2026-09-27