AI Potluck
Back to Gap Map Model components / Language-specific datasets

NaijaVoices

NaijaVoices
restricted / Overall score: 2.7

NaijaVoices is a speech dataset of about 1,800 hours of speech paired with expert-curated text in Igbo, Hausa and Yoruba, roughly 600 hours per language, from more than 5,000 speakers with a balanced age range. Each row carries the audio, transcript, language, speaker ID, gender and age bracket. The NaijaVoices community project collected it and also publishes a compressed copy of the same data.

Openness

2 high confidence
2.0
license
cc-by-nc-sa-4.0(card: commercial use needs a paid membership waiver)
access
auto(Hugging Face automatic gate
dataset_card
present(composition, fields, batches and splits described

The data is free for non-commercial research under a share-alike license, and commercial use requires one of the project's paid membership tiers. Downloading it means accepting the project's usage terms through an automatic Hugging Face gate.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the full repository and its compressed copy, which hold the same data. Downloads count file fetches behind a click-through gate, not distinct users.

Capability

3 medium confidence
3.0

NaijaVoices gives Igbo, Hausa and Yoruba about six hundred hours each from thousands of speakers with expert-curated text and a paper on its collection, and recognizers such as a Hausa model from asr-africa are trained on it. At under two thousand hours across three languages it holds less than Afrivoice, which has thousands of hours across eighteen, and far less than the hundreds of thousands in the largest English speech corpora.

Verified 2026-09-24