AI Potluck
Model components / Benchmark / eval datasets

MMMU

Unknown

Massive Multi-discipline Multimodal Understanding: about 11,550 college-level questions combining images and text across 30 subjects and 183 subfields, using 30 image types, drawn from college exams, quizzes and textbooks. It splits into a 150-question dev set, a 900-question validation set and a 10,500-question test set.

Test-set answers were held out behind an evaluation server until 12 February 2026 and are now published, so every split is locally answerable. Verified 2026-08-13 via the MMMU/MMMU dataset card on Hugging Face and the MMMU paper abstract.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI/open,redistributable)
access
public(not gated)
datasheet
present(comprehensive card)
splits
dev+validation+test all public (test answers released 2026-02-12

Fully open: Apache-2.0, ungated, redistributable, comprehensive card. Test-set answers were released 2026-02-12 (previously held out via EvalAI), so there is no longer a held-out element; scored 5 as a fully-public benchmark.

  • https://huggingface.co/datasets/MMMU/MMMU recorded 2026-08-13

    `license:apache-2.0` in the repo tags; the embedded repo state reads `"gated":false`; the card news list still reads "[2026-02-12]: We have released the answers for the test set!" and the EvalAI submission paragraph is struck through; dev/validation/test splits all present; 51,175 downloads in the trailing 30 days

  • https://arxiv.org/abs/2311.16502 recorded 2026-08-13

    the abstract still describes 11.5K multimodal questions across six core disciplines from college exams, quizzes and textbooks

Adoption

3 high confidence
3.0

51,058 Hugging Face downloads in the trailing 30 days for MMMU/MMMU, which puts it in the 10K-100K band, level 3 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. MMMU and MMMU-Pro are the standard multimodal-understanding benchmarks cited in frontier multimodal model reports.

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored; openness and adoption carry this category.

  • https://huggingface.co/datasets/MMMU/MMMU recorded 2026-08-13

    a static image+text question corpus across 30 subjects; no performance, throughput or feature claim on the page for the capability axis to read

Verified 2026-08-12