AI Potluck
Model components / Benchmark / eval datasets

MMLU-Pro

TIGER-Lab

A harder successor to MMLU with 12,000 complex multiple-choice questions across many disciplines. It is designed to benchmark large language models more rigorously than the original MMLU and has become a common upgrade path for general knowledge and reasoning evals.

MMLU-Pro, ~12,032 questions across 14 disciplines with 10 answer options (vs MMLU's 4), expert-reviewed and contamination-hardened; test split (12,000) + small validation (70). Paper arXiv:2406.01574. HF dataset TIGER-Lab/MMLU-Pro verified live June 2026.

Openness

5 high confidence
5.0
license
MIT(OSI/open,redistributable)
access
public(not gated)
datasheet
present(comprehensive card+leaderboard)
splits
public(test+validation, no hidden held-out)

Fully open: MIT, ungated, redistributable, comprehensive card.

Adoption

4 high confidence
4.0

~154,919 HF downloads in the trailing month (primary, the highest of this batch); MMLU-Pro is a de-facto-standard knowledge/reasoning benchmark cited across frontier model reports (appears as a capability basis throughout the v3 flagship set).

Capability

not assessed

Dataset, not 'capable'; openness and adoption carry this category per recipe.

Unchanged since 2026-06-09 (last edited, not re-checked)