Inkuba-Mono
Lelapa AIInkuba-Mono is a monolingual text corpus for Hausa, Yoruba, Swahili, isiZulu and isiXhosa, about 1.7 billion tokens compiled from public repositories on Hugging Face, GitHub and Zenodo. Swahili contributes about a billion tokens and Yoruba the fewest. Lelapa AI assembled it to pretrain its InkubaLM small language model.
Openness
2 medium confidence- license
- cc-by-nc-4.0(card body
- access
- auto(automatic Hugging Face gate)
- dataset_card
- present(per-language sentence and token counts
It sits behind an automatic click-through and the card restricts it to non-commercial use. The card names the kinds of repositories it drew from but not the individual source corpora or their terms.
- https://huggingface.co/api/datasets/lelapa/Inkuba-Mono recorded 2026-09-24
"gated": "auto"; no license in cardData.
- https://huggingface.co/datasets/lelapa/Inkuba-Mono recorded 2026-09-24
"We collected monolingual datasets to train our InkubaLM for 5 African languages from the following public repositories. Hugging Face Github Zenodo"; "License: CC BY-NC 4.0".
Adoption
1 high confidenceHugging Face downloads of the single Inkuba-Mono repository, behind an automatic gate. Downloads count file fetches, not users or trained models.
- https://huggingface.co/api/datasets/lelapa/Inkuba-Mono recorded 2026-09-24
56 downloads in the trailing 30 days for lelapa/Inkuba-Mono
Capability
2 medium confidenceInkuba-Mono is the pretraining text behind Lelapa AI's InkubaLM, with about a billion tokens of Swahili among its five African languages, compiled from existing public corpora that its card does not list. At under two billion tokens it is smaller than WURA, which also documents its own audit of the web text it draws on, and a tiny fraction of the trillions in English pretraining corpora.
- https://huggingface.co/datasets/lelapa/Inkuba-Mono recorded 2026-09-24
Statistics table: Hausa 359 M, Yoruba 77 M, Swahili 1B, isiZulu 207M, isiXhosa 79M tokens; models trained on it include lelapa/InkubaLM-0.4B.
Verified 2026-09-24