Kencorpus
lab · KenyaScores
1 product on the map — 1 open.
Openness
5 high confidence- license
- cc-by-4.0(card metadata on all six repositories)
- access
- public(Hugging Face repositories ungated)
- dataset_card
- present(each repository has statistics, fields and usage)
All six repositories download without a gate under an attribution-only license. The paper calls it one of few public-domain corpora for these languages, though the cards use CC BY 4.0 rather than a public-domain dedication.
- https://arxiv.org/abs/2208.12081 recorded 2026-09-24
"Kencorpus is one of few public domain corpora for these three low resource languages".
- https://huggingface.co/api/datasets/Kencorpus/KenCorpus_text recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/Kencorpus/KenCorpus_text/raw/main/README.md recorded 2026-09-24
Front matter "license: cc-by-4.0"; files per language and sources described.
Adoption
1 high confidenceHugging Face downloads summed over the text, audio and four task repositories that carry the Kencorpus paper link. Downloads count file fetches, not users.
- https://huggingface.co/api/datasets/Kencorpus/KenCorpus_audio recorded 2026-09-24
100 downloads in the trailing 30 days for Kencorpus/KenCorpus_audio
- https://huggingface.co/api/datasets/Kencorpus/KenCorpus_text recorded 2026-09-24
164 downloads in the trailing 30 days for Kencorpus/KenCorpus_text
- https://huggingface.co/api/datasets/Kencorpus/KenPOS recorded 2026-09-24
58 downloads in the trailing 30 days for Kencorpus/KenPOS
- https://huggingface.co/api/datasets/Kencorpus/KenSpeech recorded 2026-09-24
316 downloads in the trailing 30 days for Kencorpus/KenSpeech
- https://huggingface.co/api/datasets/Kencorpus/KenSwQuAD recorded 2026-09-24
132 downloads in the trailing 30 days for Kencorpus/KenSwQuAD
- https://huggingface.co/api/datasets/Kencorpus/KenTrans recorded 2026-09-24
112 downloads in the trailing 30 days for Kencorpus/KenTrans
Capability
2 medium confidenceKencorpus gives Swahili, Dholuo and Luhya a community-sourced text and speech corpus with tagging, question answering and translation sets, documented in a paper, with one community translation model trained on it. At under two hundred hours of speech and a few million words of text it is small beside English corpora of hundreds of thousands of hours and trillions of tokens, and WURA's African web corpus trained the AfriTeVa models.
- https://arxiv.org/abs/2208.12081 recorded 2026-09-24
Abstract: "5,594 items - 4,442 texts (5.6M words) and 1,152 speech files (177hrs)"; POS sets, 7,537 QA pairs and 13,400 translation sentences.
- https://huggingface.co/datasets/Kencorpus/KenSwQuAD/raw/main/README.md recorded 2026-09-24
"7,506 question-answer pairs derived from 1,441 unique Swahili contexts".