AI Potluck
Back to Gap Map Model components / Language-specific datasets

Kencorpus

Kencorpus
open / Overall score: 1.7

Kencorpus is a text and speech corpus for Swahili, Dholuo and Luhya (Lumarachi, Logooli and Lubukusu), gathered from communities, schools, media and publishers: 4,442 texts of about 5.6 million words and 177 hours of spontaneous speech. Companion sets add part-of-speech tags, Swahili question answering, translation into Swahili and a Swahili speech recognition set. The Kencorpus project publishes it.

Openness

5 high confidence
5.0
license
cc-by-4.0(card metadata on all six repositories)
access
public(Hugging Face repositories ungated)
dataset_card
present(each repository has statistics, fields and usage)

All six repositories download without a gate under an attribution-only license. The paper calls it one of few public-domain corpora for these languages, though the cards use CC BY 4.0 rather than a public-domain dedication.

Adoption

1 high confidence
1.0

Hugging Face downloads summed over the text, audio and four task repositories that carry the Kencorpus paper link. Downloads count file fetches, not users.

Capability

2 medium confidence
2.0

Kencorpus gives Swahili, Dholuo and Luhya a community-sourced text and speech corpus with tagging, question answering and translation sets, documented in a paper, with one community translation model trained on it. At under two hundred hours of speech and a few million words of text it is small beside English corpora of hundreds of thousands of hours and trillions of tokens, and WURA's African web corpus trained the AfriTeVa models.

Verified 2026-09-24