AI Potluck
Back to Gap Map Model components / Language-specific datasets

AfroBench

McGill NLP
open / Overall score: 3.8

AfroBench is an evaluation suite for large language models covering 64 African languages, 15 tasks and 22 datasets. It carries AfriMMLU, AfriXNLI, AfriMGSM, AfriSenti, MasakhaNER, MasakhaNEWS, MasakhaPOS, AfriQA, MAFAND-MT, SALT, Uhura, InjongoIntent, NaijaRC, NollySenti, NTREX African and AfriADR, with lm-evaluation-harness configs and a public leaderboard. It is published from the McGill-NLP GitHub organization, with its data in a Masakhane collection.

Belebele, FLORES, SIB-200, MMMLU and XL-Sum are in the suite but are separate entries, so their repositories are not listed or counted here. NaijaRC and NollySenti sit in a personal Hugging Face namespace and are also left out.

Openness

4 medium confidence
4.0
license
mixed-per-subset(afrimmlu: apache-2.0
afrixnli
apache-2.0
afrimgsm
apache-2.0
afrisenti
cc-by-nc-sa-2.0
masakhaner-x
unknown
masakhanews
afl-3.0
masakhapos
afl-3.0
afriqa-gold-passages
cc-by-sa-4.0
mafand
cc-by-nc-4.0
uhura-arc-easy
mit
InjongoIntent
apache-2.0
ntrex_african
cc-by-sa-4.0
AfriADR
apache-2.0
salt
cc-by-sa-4.0
access
public(Masakhane member repositories ungated
dataset_card
present(README, project site and paper describe tasks and datasets

Each member dataset keeps its own license, from Apache-2.0 and MIT to non-commercial terms on MAFAND-MT and AfriSenti, and MasakhaNER-X lists its license as unknown. Everything but SALT downloads without a gate.

Adoption

1 low confidence
1.0

GitHub stars on the AfroBench repository. The member datasets it evaluates are also downloaded on their own, but those downloads are not runs of the suite, so they are left out; a star is attention rather than use.

Capability

5 medium confidence
5.0

AfroBench brings 22 datasets across 64 African languages into one evaluation with harness tasks and a public leaderboard, and models such as InkubaLM report on its members. No other open suite here covers as many African languages and tasks. Its breadth comes from bundling existing Masakhane and partner benchmarks.

Verified 2026-09-24