AI Potluck
Back to Gap Map Model components / Language-specific datasets

Afrivoice

Digital Umuganda
gated / Overall score: 2.7

Afrivoice is a family of speech datasets, mostly recorded by speakers describing photographs aloud, with transcriptions for all or part of the audio. Releases cover Kinyarwanda, Swahili, five Ethiopian languages, a set including Shona, Lingala, Wolof and Somali, and a batch for Kirundi, Ndau, Ndebele, Oshiwambo and Gamo, about 15,000 hours across the current releases. Digital Umuganda collects and publishes it.

The artifact list includes superseded and legacy repositories, which are counted with the current releases.

Openness

3 medium confidence
3.0
license
cc-by-4.0(card metadata on eight repositories
access
auto(eight repositories behind an automatic Hugging Face gate
dataset_card
present(current releases describe domains, hours and fields

Most repositories open after an automatic click-through on Hugging Face, and their metadata carries an attribution-only license. The Swahili instruction-format copy has neither a license nor a card.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over all nine Afrivoice repositories, most of them behind an automatic gate. Downloads count file fetches; audio reused inside WAXAL is counted there, not here.

Capability

3 medium confidence
3.0

Afrivoice gives eighteen African languages image-described and scripted speech across five releases, more than any other open collection for most of them, and Digital Umuganda's Mbaza recognizer is trained on it while Google's WAXAL corpus draws on it, though it has release cards but no paper. Some releases transcribe only part of the audio, leaving a few thousand transcribed hours, far from the hundreds of thousands in the largest English speech corpora.

Verified 2026-09-24