KazMMLU
Mohamed bin Zayed University of Artificial IntelligenceKazMMLU tests language models on 23,000 multiple-choice questions in Kazakh and Russian, 10,969 and 12,031 respectively, reflecting Kazakhstan's bilingual education system. The questions cover STEM, humanities, and social sciences in high-school and university material, sourced from authentic educational materials and manually validated by native speakers and educators. MBZUAI researchers built it.
The Hub card is a stub from before the repository was renamed and does not describe the release; the description follows the paper.
Openness
2 high confidence- license
- cc-by-nc-4.0(card metadata and API tag)
- access
- public(ungated on the Hub)
- dataset_card
- partial(card lists splits and columns but is titled with another repository name and truncates its subset list)
The questions download freely under a noncommercial license, which rules out commercial evaluation use. The card covers only the file layout, so how the questions were sourced is documented in the paper alone.
- https://huggingface.co/api/datasets/MBZUAI/KazMMLU recorded 2026-09-24
API JSON: "gated": false; tag license:cc-by-nc-4.0.
- https://huggingface.co/datasets/MBZUAI/KazMMLU/raw/main/README.md recorded 2026-09-24
Card metadata "license: cc-by-nc-4.0"; body titled "# MukhammedTogmanov/jana"; subset list ends "... (list other subsets)".
Adoption
1 high confidenceHugging Face downloads of the Hub repository. Evaluation harnesses fetch a benchmark on each run, so the figure counts runs rather than groups.
- https://huggingface.co/api/datasets/MBZUAI/KazMMLU recorded 2026-09-24
392 downloads in the trailing 30 days for MBZUAI/KazMMLU
Capability
3 medium confidenceKazMMLU is the first MMLU-style benchmark for Kazakh, built from native educational material and documented in a paper. No leaderboard or model report outside that paper was found to use it, and its card is a placeholder, which places it a step below ArabicMMLU, from the same lab.
- https://arxiv.org/abs/2502.12829 recorded 2026-09-24
Abstract: "KazMMLU, the first MMLU-style dataset specifically designed for Kazakh language. KazMMLU comprises 23,000 questions ... sourced from authentic educational materials and manually validated by native speakers and educators".
- https://huggingface.co/datasets/MBZUAI/KazMMLU recorded 2026-09-24
Hub page: "Number of rows: 23,000"; no models or Spaces listed as using it.
Verified 2026-09-24