AI Potluck
Model components / Benchmark / eval datasets

MMMU

Unknown

MMMU (Massive Multi-discipline Multimodal Understanding), ~11,550 college-level multimodal (image+text) questions across 30 subjects/183 subfields, 30 image types. Splits: dev 150 / validation 900 / test 10,500. Test-set answers were previously held out behind EvalAI but were RELEASED on 2026-02-12, so all splits are now publicly answerable. CVPR 2024, paper arXiv:2311.16502. HF dataset MMMU/MMMU verified live June 2026.

MMMU (Massive Multi-discipline Multimodal Understanding), ~11,550 college-level multimodal (image+text) questions across 30 subjects/183 subfields, 30 image types. Splits: dev 150 / validation 900 / test 10,500. Test-set answers were previously held out behind EvalAI but were RELEASED on 2026-02-12, so all splits are now publicly answerable. CVPR 2024, paper arXiv:2311.16502. HF dataset MMMU/MMMU verified live June 2026.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI/open,redistributable)
access
public(not gated)
datasheet
present(comprehensive card)
splits
dev+validation+test all public (test answers released 2026-02-12

Fully open: Apache-2.0, ungated, redistributable, comprehensive card. Test-set answers were released 2026-02-12 (previously held out via EvalAI), so there is no longer a held-out element; scored 5 as a fully-public benchmark.

Adoption

4 high confidence
4.0

~86,107 HF downloads in the trailing month (primary); MMMU / MMMU-Pro is the de-facto-standard multimodal-understanding benchmark cited in frontier multimodal model reports (e.g. GPT-5 reports MMMU; Gemini reports MMMU-Pro in the v3 flagship set).

Capability

not assessed

Dataset, not 'capable'; openness and adoption carry this category per recipe.

Unchanged since 2026-06-09 (last edited, not re-checked)