MMMU
UnknownMMMU (Massive Multi-discipline Multimodal Understanding), ~11,550 college-level multimodal (image+text) questions across 30 subjects/183 subfields, 30 image types. Splits: dev 150 / validation 900 / test 10,500. Test-set answers were previously held out behind EvalAI but were RELEASED on 2026-02-12, so all splits are now publicly answerable. CVPR 2024, paper arXiv:2311.16502. HF dataset MMMU/MMMU verified live June 2026.
MMMU (Massive Multi-discipline Multimodal Understanding), ~11,550 college-level multimodal (image+text) questions across 30 subjects/183 subfields, 30 image types. Splits: dev 150 / validation 900 / test 10,500. Test-set answers were previously held out behind EvalAI but were RELEASED on 2026-02-12, so all splits are now publicly answerable. CVPR 2024, paper arXiv:2311.16502. HF dataset MMMU/MMMU verified live June 2026.
Openness
5 high confidence- license
- Apache-2.0(OSI/open,redistributable)
- access
- public(not gated)
- datasheet
- present(comprehensive card)
- splits
- dev+validation+test all public (test answers released 2026-02-12
Fully open: Apache-2.0, ungated, redistributable, comprehensive card. Test-set answers were released 2026-02-12 (previously held out via EvalAI), so there is no longer a held-out element; scored 5 as a fully-public benchmark.
- https://huggingface.co/datasets/MMMU/MMMU recorded 2026-06-04
Apache-2.0, not gated, 11,550 questions (dev 150/val 900/test 10,500); test answers released 2026-02-12 (no longer EvalAI-held-out); 86,107 downloads last month
- https://arxiv.org/abs/2311.16502 recorded 2026-06-04
MMMU paper documenting 30 subjects, 183 subfields, multimodal benchmark
Adoption
4 high confidence~86,107 HF downloads in the trailing month (primary); MMMU / MMMU-Pro is the de-facto-standard multimodal-understanding benchmark cited in frontier multimodal model reports (e.g. GPT-5 reports MMMU; Gemini reports MMMU-Pro in the v3 flagship set).
- https://huggingface.co/datasets/MMMU/MMMU recorded 2026-06-04
86,107 downloads last month; EvalAI evaluation infrastructure
Capability
not assessedDataset, not 'capable'; openness and adoption carry this category per recipe.
Unchanged since 2026-06-09 (last edited, not re-checked)