NaijaVoices
companyScores
1 product on the map — 1 closed.
Openness
2 high confidence- license
- cc-by-nc-sa-4.0(card: commercial use needs a paid membership waiver)
- access
- auto(Hugging Face automatic gate
- dataset_card
- present(composition, fields, batches and splits described
The data is free for non-commercial research under a share-alike license, and commercial use requires one of the project's paid membership tiers. Downloading it means accepting the project's usage terms through an automatic Hugging Face gate.
- https://huggingface.co/api/datasets/naijavoices/naijavoices-dataset recorded 2026-09-24
"gated": "auto"; cardData license "cc-by-nc-sa-4.0".
- https://huggingface.co/datasets/naijavoices/naijavoices-dataset recorded 2026-09-24
"The NaijaVoices dataset is licensed under the CC BY-NC-SA 4.0 license... For commercial interests, please check out our membership tiers which offer special commercial waivers."
Adoption
2 high confidenceHugging Face downloads summed over the full repository and its compressed copy, which hold the same data. Downloads count file fetches behind a click-through gate, not distinct users.
- https://huggingface.co/api/datasets/naijavoices/naijavoices-dataset recorded 2026-09-24
1130 downloads in the trailing 30 days for naijavoices/naijavoices-dataset
- https://huggingface.co/api/datasets/naijavoices/naijavoices-dataset-compressed recorded 2026-09-24
54 downloads in the trailing 30 days for naijavoices/naijavoices-dataset-compressed
Capability
3 medium confidenceNaijaVoices gives Igbo, Hausa and Yoruba about six hundred hours each from thousands of speakers with expert-curated text and a paper on its collection, and recognizers such as a Hausa model from asr-africa are trained on it. At under two thousand hours across three languages it holds less than Afrivoice, which has thousands of hours across eighteen, and far less than the hundreds of thousands in the largest English speech corpora.
- https://arxiv.org/abs/2505.20564 recorded 2026-09-24
Abstract: "a 1,800-hour speech-text dataset with 5,000+ speakers" with ASR fine-tuning results for Whisper, MMS and XLSR.
- https://huggingface.co/datasets/naijavoices/naijavoices-dataset recorded 2026-09-24
Card: "1,800 hours of authentic speech (from over 5,000 diverse speakers!)... ~600 hours for each of the three languages"; models trained on it include asr-africa/w2v-bert-2.0-naijavoices-hausa-500hr-v0.