AI Potluck
Back to Gap Map Model components / Language-specific datasets

Inkuba-Instruct

Lelapa AI
restricted / Overall score: 3.1

Inkuba-Instruct is an instruction-tuning dataset for Hausa, Yoruba, Swahili, isiZulu and isiXhosa, built by recasting existing task datasets such as MAFAND-MT, MasakhaNER, MasakhaPOS, AfriQA, SIB-200, MasakhaNEWS and AfriSenti into instruction, input and output records. It spans translation, named-entity recognition, part-of-speech tagging, question answering, topic classification and sentiment. Lelapa AI released it with the InkubaLM paper.

Only the train and dev splits are public; the card says the test split will follow later.

Openness

2 medium confidence
2.0
license
cc-by-nc-4.0(card body
access
auto(automatic Hugging Face gate
dataset_card
present(source datasets per task, sample counts and record structure)

The card restricts it to non-commercial use behind an automatic click-through, and its test split is not yet released. The source datasets keep their own licenses, which the card does not reconcile with its own.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single repository, behind an automatic gate. Downloads count file fetches, not models tuned on it.

Capability

4 medium confidence
4.0

Inkuba-Instruct turns well-known African task datasets into instruction data for five languages, lists every source on its card, and the Pula-8B model is tuned on it. Like WangchanThaiInstruct it is documented and has a named model tuned on it, but its content recasts existing human-annotated sets rather than new writing by people.

Verified 2026-09-24