AI Potluck
Back to Gap Map Model components / Language-specific datasets

Inkuba-Mono

Lelapa AI
restricted / Overall score: 1.7

Inkuba-Mono is a monolingual text corpus for Hausa, Yoruba, Swahili, isiZulu and isiXhosa, about 1.7 billion tokens compiled from public repositories on Hugging Face, GitHub and Zenodo. Swahili contributes about a billion tokens and Yoruba the fewest. Lelapa AI assembled it to pretrain its InkubaLM small language model.

Openness

2 medium confidence
2.0
license
cc-by-nc-4.0(card body
access
auto(automatic Hugging Face gate)
dataset_card
present(per-language sentence and token counts

It sits behind an automatic click-through and the card restricts it to non-commercial use. The card names the kinds of repositories it drew from but not the individual source corpora or their terms.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single Inkuba-Mono repository, behind an automatic gate. Downloads count file fetches, not users or trained models.

Capability

2 medium confidence
2.0

Inkuba-Mono is the pretraining text behind Lelapa AI's InkubaLM, with about a billion tokens of Swahili among its five African languages, compiled from existing public corpora that its card does not list. At under two billion tokens it is smaller than WURA, which also documents its own audit of the web text it draws on, and a tiny fraction of the trillions in English pretraining corpora.

Verified 2026-09-24