MMS-lab data
MetaMMS-lab is Meta's labeled speech corpus of New Testament readings in 1,107 languages, built for the Massively Multilingual Speech project. Recordings from Faith Comes By Hearing, goto.bible and bible.com were force-aligned with their texts chapter by chapter and filtered for quality, and nearly every recording has a single speaker. Meta trained its MMS recognition and synthesis models on it.
Openness
1 high confidence- license
- not-clearly-stated-on-card(Not distributed
- access
- closed(The paper releases models, not the aligned data.)
- dataset_card
- no(No card
Meta released the models trained on this data but not the aligned recordings, which come from third-party Bible audio publishers.
- https://arxiv.org/abs/2305.13516 recorded 2026-09-24
"a new dataset based on readings of publicly available religious texts".
- https://arxiv.org/pdf/2305.13516 recorded 2026-09-24
"The MMS models are available at https://github.com/pytorch/fairseq/tree/master/examples/mms"; data drawn from "Faith Comes By Hearing, goto.bible and bible.com"; no data download is offered.
Adoption
not assessedThe aligned data was never distributed, so there is no download or usage figure to read.
- https://arxiv.org/abs/2305.13516 recorded 2026-09-24
The MMS paper releases the models trained on the data, not the aligned recordings, so there is no download or user figure to read
Capability
4 medium confidenceMMS-lab spans more languages than any other labeled speech corpus, and Meta trained its MMS recognition and synthesis models for more than a thousand languages on it. It is a step below Mozilla Common Voice because it is one narrow domain, Bible reading, mostly by a single speaker per recording, and the MMS paper found models trained on Common Voice did better on FLEURS.
- https://arxiv.org/abs/2305.13516 recorded 2026-09-24
"a single multilingual automatic speech recognition model for 1,107 languages, speech synthesis models for the same number of languages".
- https://arxiv.org/pdf/2305.13516 recorded 2026-09-24
"speech audio paired with corresponding text in 1,107 languages (MMS-lab; 44.7K hours)"; "it is both from a particular narrow domain and most recordings are from a single speaker".
Verified 2026-09-24