BibleTTS
MasakhaneBibleTTS is a text-to-speech corpus of studio-quality, single-speaker Bible readings at 48kHz, with up to 86 hours of aligned speech per language. OpenSLR hosts aligned audio and text for six languages, Akuapem Twi, Asante Twi, Ewe, Hausa, Lingala and Yoruba, and the paper covers four more without alignment. Researchers from Masakhane and Coqui built it from Biblica's open.bible recordings.
There is no Hugging Face copy; the data is distributed through OpenSLR resource 129.
Openness
5 high confidence- license
- cc-by-sa-4.0(OpenSLR resource page
- access
- public(direct .tgz downloads from OpenSLR mirrors)
- dataset_card
- present(OpenSLR page and paper describe contents and alignment)
The six aligned languages download directly from OpenSLR under a share-alike license that allows commercial use. The recordings are a derivative of Biblica's open.bible project rather than new recordings.
- https://arxiv.org/abs/2207.03546 recorded 2026-09-24
Abstract: "a derivative work of Bible recordings made and released by the open.bible project from Biblica"; "released under a commercial-friendly CC-BY-SA license".
- https://raw.githubusercontent.com/masakhane-io/bibleTTS/gh-pages/README.md recorded 2026-09-24
README links the project website, paper, alignment code and Coqui TTS training; no license stated.
- https://www.openslr.org/129/ recorded 2026-09-24
License: "Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)"; downloads akuapem-twi.tgz, asante-twi.tgz, ewe.tgz, hausa.tgz, lingala.tgz, yoruba.tgz.
Adoption
1 low confidenceGitHub stars on the BibleTTS repository; the audio itself is downloaded from OpenSLR, which publishes no count, and a star is attention rather than use.
- https://ungh.cc/repos/masakhane-io/bibleTTS recorded 2026-09-24
GitHub repository record for masakhane-io/bibleTTS (via the ungh.cc mirror of the GitHub API): 14 stargazers, last push 2022-09-20.
Capability
2 medium confidenceBibleTTS provides studio-quality single-speaker Bible readings for six African languages, with a paper on alignment and speech synthesis baselines, though no model beyond that paper uses it. At a few hundred hours of Bible text it is small beside English speech corpora of hundreds of thousands of hours, and Afrivoice records more varied speech across eighteen languages, Lingala among them.
- https://arxiv.org/abs/2207.03546 recorded 2026-09-24
Abstract: "up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language"; results for TTS models with Coqui TTS.
Verified 2026-09-24