AI Potluck
Back to Gap Map Model components / Language-specific datasets

AfriVoices-KE

African Next Voices Kenya
gated / Overall score: 2.7

AfriVoices-KE is a speech dataset of about 3,000 hours in five Kenyan languages, Dholuo, Kikuyu, Kalenjin, Maasai and Somali, recorded on smartphones by 4,777 native speakers. About 750 hours are scripted readings across eleven domains and the rest is spontaneous speech prompted by text and images, checked by human reviewers. The Kenyan arm of the African Next Voices project publishes one repository per language.

The Hugging Face repositories carry no dataset card, so the license and composition are taken from the paper.

Openness

3 medium confidence
3.0
license
cc-by-4.0(stated in the paper, Section 7
access
manual(users must share contact information and accept conditions on Hugging Face)
dataset_card
no("No dataset card yet" on all five repositories

The paper releases the audio under an attribution-only license, but each language repository asks for contact details and acceptance of conditions before download. The repositories themselves say nothing about license or composition.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the five per-language repositories, each behind a gate that collects contact details. Downloads count file fetches, not approved users.

Capability

3 medium confidence
3.0

AfriVoices-KE gives five Kenyan languages scripted and spontaneous speech from thousands of native speakers with a detailed paper, and recognizers such as an East African model and a Kalenjin Whisper model are trained on it, though its repositories lack a full dataset card. Its roughly three thousand hours match its South African sibling, ZA African Next Voices, and are a small share of the hundreds of thousands in the largest English speech corpora.

  • https://arxiv.org/abs/2604.08448 recorded 2026-09-24

    Abstract: "approximately 3,000 hours of audio across five Kenyan languages... 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers".

  • https://huggingface.co/datasets/Anv-ke/Kalenjin recorded 2026-09-24

    "Models trained or fine-tuned on Anv-ke/Kalenjin" lists leophill/w2v-bert-2.0-afrivoices-ea-6l-asr and Rekody/whisper-large-v3-turbo-kalenjin.

Verified 2026-09-24