AI2 ARC
Allen Institute for AIBenchmark of 7,787 grade-school science questions assembled to study advanced question answering, split into an Easy set of 5,197 and a Challenge set of 2,590. Published by the Allen Institute for AI in 2018, it remains a standard multiple-choice reasoning evaluation in model release reports.
Verified 2026-08-13 via the allenai/ai2_arc dataset card on Hugging Face.
Openness
5 high confidence- data
- open(7,787 questions, downloadable on HF)
- license
- CC-BY-SA-4.0(open data license, redistributable)
- datasheet
- yes(HF dataset card + arXiv paper)
- access
- public(train/val/test splits all released, not held-out)
Fully open under CC-BY-SA-4.0 with dataset card; the canonical open-data benchmark profile.
- https://huggingface.co/datasets/allenai/ai2_arc recorded 2026-08-13
`license:cc-by-sa-4.0` in the repo tags; the embedded repo state reads `"gated":false`, so train/validation/test all download without a barrier; the dataset card still renders; 474,701 downloads in the trailing 30 days
Adoption
4 high confidence472,342 Hugging Face downloads in the trailing 30 days for allenai/ai2_arc, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. ARC is one of the most-cited standard reasoning benchmarks, routinely reported in frontier model release cards.
- https://huggingface.co/api/datasets/allenai/ai2_arc recorded 2026-08-12
472,342 downloads in the trailing 30 days for allenai/ai2_arc
Capability
not assessedCapability is not a meaningful axis for an eval dataset, so it is left unscored.
- https://huggingface.co/datasets/allenai/ai2_arc recorded 2026-08-13
a static grade-school science question corpus; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12