AI Potluck
Back to Gap Map Model components / Language-specific datasets

BibleTTS

Masakhane
open / Overall score: 1.7

BibleTTS is a text-to-speech corpus of studio-quality, single-speaker Bible readings at 48kHz, with up to 86 hours of aligned speech per language. OpenSLR hosts aligned audio and text for six languages, Akuapem Twi, Asante Twi, Ewe, Hausa, Lingala and Yoruba, and the paper covers four more without alignment. Researchers from Masakhane and Coqui built it from Biblica's open.bible recordings.

There is no Hugging Face copy; the data is distributed through OpenSLR resource 129.

Openness

5 high confidence
5.0
license
cc-by-sa-4.0(OpenSLR resource page
access
public(direct .tgz downloads from OpenSLR mirrors)
dataset_card
present(OpenSLR page and paper describe contents and alignment)

The six aligned languages download directly from OpenSLR under a share-alike license that allows commercial use. The recordings are a derivative of Biblica's open.bible project rather than new recordings.

Adoption

1 low confidence
1.0

GitHub stars on the BibleTTS repository; the audio itself is downloaded from OpenSLR, which publishes no count, and a star is attention rather than use.

Capability

2 medium confidence
2.0

BibleTTS provides studio-quality single-speaker Bible readings for six African languages, with a paper on alignment and speech synthesis baselines, though no model beyond that paper uses it. At a few hundred hours of Bible text it is small beside English speech corpora of hundreds of thousands of hours, and Afrivoice records more varied speech across eighteen languages, Lingala among them.

  • https://arxiv.org/abs/2207.03546 recorded 2026-09-24

    Abstract: "up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language"; results for TTS models with Coqui TTS.

Verified 2026-09-24