AI Potluck
Model components / Benchmark / eval datasets

MMLU-Pro

TIGER-Lab

Harder successor to MMLU with about 12,032 questions across 14 disciplines, expanding the choice set from four options to ten and removing trivial and noisy items. It is expert-reviewed and contamination-hardened, and ships a 12,000-question test split with a small validation set.

Verified 2026-08-13 via the TIGER-Lab/MMLU-Pro dataset card on Hugging Face and the MMLU-Pro paper abstract.

Openness

5 high confidence
5.0
license
MIT(OSI/open,redistributable)
access
public(not gated)
datasheet
present(comprehensive card+leaderboard)
splits
public(test+validation, no hidden held-out)

Fully open: MIT, ungated, redistributable, comprehensive card.

  • https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro recorded 2026-08-13

    `license:mit` in the repo tags; the embedded repo state reads `"gated":false`; the dataset card and leaderboard link still render; 172,078 downloads in the trailing 30 days

  • https://arxiv.org/abs/2406.01574 recorded 2026-08-13

    the abstract still describes MMLU-Pro as expanding the choice set "from four to ten options" and eliminating trivial and noisy MMLU questions

Adoption

4 high confidence
4.0

170,681 Hugging Face downloads in the trailing 30 days for TIGER-Lab/MMLU-Pro, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. MMLU-Pro is a standard knowledge-and-reasoning benchmark cited across frontier model reports.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored; openness and adoption carry this category.

Verified 2026-08-12