AI Potluck
Model components / Benchmark / eval datasets

MMLU

Hendrycks et al. (UC Berkeley)

Massive Multitask Language Understanding: 14,042 four-option multiple-choice test questions across 57 subjects spanning humanities, social sciences, STEM and professional domains, designed to measure broad knowledge and reasoning. An auxiliary training split brings the repository to roughly 231,400 rows.

Largely superseded by MMLU-Pro for separating frontier models. Verified 2026-08-13 via the cais/mmlu dataset card on Hugging Face.

Openness

5 high confidence
5.0
data
downloadable(HF parquet, cais/mmlu)
license
MIT(OSI/open)
datasheet
present(card documents 57 subjects, splits dev/val/test, schema)
redistributable
yes
paper
public(ICLR 2021)

Clean open data: MIT-licensed, downloadable, with a documented datasheet and a public paper. Fully open.

  • https://huggingface.co/datasets/cais/mmlu recorded 2026-08-13

    `license:mit` in the repo tags; the embedded repo state reads `"gated":false`, so the parquet downloads without a barrier; the dataset card documents the 57 subjects and the dev/val/test splits; 477,123 downloads in the trailing 30 days

Adoption

4 high confidence
4.0

477,890 Hugging Face downloads in the trailing 30 days for cais/mmlu, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. MMLU is the canonical general-knowledge LLM benchmark, reported in essentially every frontier model release and still a universal reference even after saturation.

Capability

not assessed

A dataset is not capability-scored, so this axis is left unscored. MMLU is saturated and contaminated - top models land between 88% and 94%, and MMLU-Pro has largely superseded it for differentiation - but that is a quality caveat rather than an openness or capability deduction.

  • https://huggingface.co/datasets/cais/mmlu recorded 2026-08-13

    a static multiple-choice question corpus across 57 subjects; no performance, throughput or feature claim on the page for the capability axis to read

Verified 2026-08-12