AI Potluck
Back to Gap Map Model components / Speech & audio

MMS (Massively Multilingual Speech)

Meta
restricted / Overall score: 2.4

Meta's Massively Multilingual Speech models: one wav2vec 2.0 recognizer with per-language adapters for over 1,100 languages, language identification across about 4,000, and a separate text-to-speech voice for each of more than 1,100 languages.

Meta's later Omnilingual ASR (2025, Apache-2.0) is a separately named system and is not folded into this entry.

Openness

2 high confidence
2.0
weights
open(ASR, language-ID and per-language VITS TTS checkpoints on the Hub, ungated)
data
closed(the MMS-lab and unlabeled corpora are not released)
code
partial(inference and adapter fine-tuning
license
CC-BY-NC-4.0(every MMS checkpoint, code and weights alike)

Every MMS checkpoint downloads freely but is licensed CC-BY-NC-4.0, which rules out commercial use. The labeled and unlabeled corpora it was trained on are not released.

Adoption

3 medium confidence
3.0

Hugging Face downloads of the six declared checkpoints. More than a thousand per-language TTS voices sit outside that figure, so the family as a whole is used more than it shows.

Capability

2 low confidence
2.0

Unmatched breadth for low-resource languages, with recognition, identification and synthesis in one family. Its accuracy is not compared on any public board, so it sits a step below Whisper, which is.

Verified 2026-09-27