AI Potluck
Back to Gap Map Model components / Language-specific datasets

ArabicMMLU

Mohamed bin Zayed University of Artificial Intelligence
restricted / Overall score: 4.1(strong)

ArabicMMLU tests language models on 14,575 multiple-choice questions in Modern Standard Arabic across 40 tasks, drawn from school exams in countries of North Africa, the Levant, and the Gulf. Ten native speakers from several countries built it, grouping questions into STEM, social science, humanities, Arabic language, and other subjects. MBZUAI leads the work with university and industry partners.

Openness

2 medium confidence
2.0
license
cc-by-nc-4.0(Hub card metadata
access
public(ungated on the Hub
dataset_card
present(card and README describe sources, construction and statistics)

Anyone can download the questions and the evaluation code, and the card and paper explain how they were collected. The noncommercial license keeps it out of commercial evaluation pipelines, and the Hub and GitHub copies disagree on whether a share-alike condition applies.

Adoption

2 high confidence
2.0

Hugging Face downloads of the Hub copy. Evaluation harnesses often fetch a benchmark on every run, so the figure counts evaluation runs as much as users, and it misses copies taken from the GitHub data folder.

Capability

5 medium confidence
5.0

ArabicMMLU was the first multi-task knowledge benchmark for Arabic, built from school exams across North Africa, the Levant and the Gulf rather than translated tests, and DarijaMMLU reuses its questions. It covers Modern Standard Arabic only, so dialect evaluation needs a companion set.

Verified 2026-09-24