ArabicMMLU
Mohamed bin Zayed University of Artificial IntelligenceArabicMMLU tests language models on 14,575 multiple-choice questions in Modern Standard Arabic across 40 tasks, drawn from school exams in countries of North Africa, the Levant, and the Gulf. Ten native speakers from several countries built it, grouping questions into STEM, social science, humanities, Arabic language, and other subjects. MBZUAI leads the work with university and industry partners.
Openness
2 medium confidence- license
- cc-by-nc-4.0(Hub card metadata
- access
- public(ungated on the Hub
- dataset_card
- present(card and README describe sources, construction and statistics)
Anyone can download the questions and the evaluation code, and the card and paper explain how they were collected. The noncommercial license keeps it out of commercial evaluation pipelines, and the Hub and GitHub copies disagree on whether a share-alike condition applies.
- https://huggingface.co/api/datasets/MBZUAI/ArabicMMLU recorded 2026-09-24
API JSON: "gated": false, "private": false.
- https://huggingface.co/datasets/MBZUAI/ArabicMMLU/raw/main/README.md recorded 2026-09-24
Card metadata: "license: cc-by-nc-4.0".
- https://raw.githubusercontent.com/mbzuai-nlp/ArabicMMLU/main/README.md recorded 2026-09-24
"The ArabicMMLU dataset is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License."
Adoption
2 high confidenceHugging Face downloads of the Hub copy. Evaluation harnesses often fetch a benchmark on every run, so the figure counts evaluation runs as much as users, and it misses copies taken from the GitHub data folder.
- https://huggingface.co/api/datasets/MBZUAI/ArabicMMLU recorded 2026-09-24
3727 downloads in the trailing 30 days for MBZUAI/ArabicMMLU
Capability
5 medium confidenceArabicMMLU was the first multi-task knowledge benchmark for Arabic, built from school exams across North Africa, the Levant and the Gulf rather than translated tests, and DarijaMMLU reuses its questions. It covers Modern Standard Arabic only, so dialect evaluation needs a companion set.
- https://arxiv.org/abs/2402.12840 recorded 2026-09-24
Abstract: "the first multi-task language understanding benchmark for the Arabic language, sourced from school exams ... 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA)".
- https://huggingface.co/datasets/MBZUAI-Paris/DarijaMMLU/raw/main/README.md recorded 2026-09-24
"translated from selected subsets of the Massive Multitask Language Understanding (MMLU) and ArabicMMLU benchmarks".
- https://raw.githubusercontent.com/mbzuai-nlp/ArabicMMLU/main/README.md recorded 2026-09-24
"The data construction process involved a total of 10 Arabic native speakers from different countries".
Verified 2026-09-24