AI Potluck
Model components / Benchmark / eval datasets

AI2 ARC

Allen Institute for AI

A benchmark of 7,787 genuine grade-school science questions assembled to study advanced question answering. The dataset is split into an Easy set and a Challenge set, and it remains one of the canonical multiple-choice reasoning benchmarks for language models.

AI2 Reasoning Challenge (ARC), Allen Institute for AI (Clark et al. 2018, arXiv:1803.05457). 7,787 grade-school science multiple-choice questions: ARC-Challenge (2,590) + ARC-Easy (5,197). Foundational LLM commonsense/reasoning benchmark, still a standard in 2026 model release reports. Verified live June 2026 on huggingface.co/datasets/allenai/ai2_arc.

Openness

5 high confidence
5.0
data
open(7,787 questions, downloadable on HF)
license
CC-BY-SA-4.0(open data license, redistributable)
datasheet
yes(HF dataset card + arXiv paper)
access
public(train/val/test splits all released, not held-out)

Fully open under CC-BY-SA-4.0 with dataset card; the canonical open-data benchmark profile.

Adoption

4 high confidence
4.0

~479K HF downloads in the last month alone (so >1M cumulative) and one of the most-cited standard reasoning benchmarks routinely reported in frontier model release cards. Strong usage_volume signal as an entrenched evaluation standard.

Capability

not assessed

Capability not a meaningful axis for an eval dataset.

Unchanged since 2026-06-09 (last edited, not re-checked)