AfriVoices-KE
African Next Voices KenyaAfriVoices-KE is a speech dataset of about 3,000 hours in five Kenyan languages, Dholuo, Kikuyu, Kalenjin, Maasai and Somali, recorded on smartphones by 4,777 native speakers. About 750 hours are scripted readings across eleven domains and the rest is spontaneous speech prompted by text and images, checked by human reviewers. The Kenyan arm of the African Next Voices project publishes one repository per language.
The Hugging Face repositories carry no dataset card, so the license and composition are taken from the paper.
Openness
3 medium confidence- license
- cc-by-4.0(stated in the paper, Section 7
- access
- manual(users must share contact information and accept conditions on Hugging Face)
- dataset_card
- no("No dataset card yet" on all five repositories
The paper releases the audio under an attribution-only license, but each language repository asks for contact details and acceptance of conditions before download. The repositories themselves say nothing about license or composition.
- https://arxiv.org/html/2604.08448 recorded 2026-09-24
"The AfriVoices-KE dataset is publicly available on Hugging Face under a CC BY 4.0 license, with access managed through a request form to track usage."
- https://huggingface.co/api/datasets/Anv-ke/Dholuo recorded 2026-09-24
"gated": "manual"; no cardData license.
- https://huggingface.co/datasets/Anv-ke/Dholuo recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "No dataset card yet".
Adoption
2 high confidenceHugging Face downloads summed over the five per-language repositories, each behind a gate that collects contact details. Downloads count file fetches, not approved users.
- https://huggingface.co/api/datasets/Anv-ke/Dholuo recorded 2026-09-24
844 downloads in the trailing 30 days for Anv-ke/Dholuo
- https://huggingface.co/api/datasets/Anv-ke/Kalenjin recorded 2026-09-24
267 downloads in the trailing 30 days for Anv-ke/Kalenjin
- https://huggingface.co/api/datasets/Anv-ke/kikuyu recorded 2026-09-24
599 downloads in the trailing 30 days for Anv-ke/kikuyu
- https://huggingface.co/api/datasets/Anv-ke/Maasai recorded 2026-09-24
248 downloads in the trailing 30 days for Anv-ke/Maasai
- https://huggingface.co/api/datasets/Anv-ke/Somali recorded 2026-09-24
198 downloads in the trailing 30 days for Anv-ke/Somali
Capability
3 medium confidenceAfriVoices-KE gives five Kenyan languages scripted and spontaneous speech from thousands of native speakers with a detailed paper, and recognizers such as an East African model and a Kalenjin Whisper model are trained on it, though its repositories lack a full dataset card. Its roughly three thousand hours match its South African sibling, ZA African Next Voices, and are a small share of the hundreds of thousands in the largest English speech corpora.
- https://arxiv.org/abs/2604.08448 recorded 2026-09-24
Abstract: "approximately 3,000 hours of audio across five Kenyan languages... 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers".
- https://huggingface.co/datasets/Anv-ke/Kalenjin recorded 2026-09-24
"Models trained or fine-tuned on Anv-ke/Kalenjin" lists leophill/w2v-bert-2.0-afrivoices-ea-6l-asr and Rekody/whisper-large-v3-turbo-kalenjin.
Verified 2026-09-24