AI Potluck
Model components / Benchmark / eval datasets

AI2 ARC

Allen Institute for AI

Benchmark of 7,787 grade-school science questions assembled to study advanced question answering, split into an Easy set of 5,197 and a Challenge set of 2,590. Published by the Allen Institute for AI in 2018, it remains a standard multiple-choice reasoning evaluation in model release reports.

Verified 2026-08-13 via the allenai/ai2_arc dataset card on Hugging Face.

Openness

5 high confidence
5.0
data
open(7,787 questions, downloadable on HF)
license
CC-BY-SA-4.0(open data license, redistributable)
datasheet
yes(HF dataset card + arXiv paper)
access
public(train/val/test splits all released, not held-out)

Fully open under CC-BY-SA-4.0 with dataset card; the canonical open-data benchmark profile.

  • https://huggingface.co/datasets/allenai/ai2_arc recorded 2026-08-13

    `license:cc-by-sa-4.0` in the repo tags; the embedded repo state reads `"gated":false`, so train/validation/test all download without a barrier; the dataset card still renders; 474,701 downloads in the trailing 30 days

Adoption

4 high confidence
4.0

472,342 Hugging Face downloads in the trailing 30 days for allenai/ai2_arc, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. ARC is one of the most-cited standard reasoning benchmarks, routinely reported in frontier model release cards.

Capability

not assessed

Capability is not a meaningful axis for an eval dataset, so it is left unscored.

Verified 2026-08-12