AI Potluck
Model components / Benchmark / eval datasets

HellaSwag

Allen Institute for AI

Commonsense natural-language-inference and sentence-completion benchmark from 2019, holding 59,950 four-way multiple-choice examples over contexts derived from ActivityNet. It tests whether a model picks the plausible continuation of an everyday situation, and is reported in most language-model releases.

The registry attributes the dataset to the Allen Institute; the primary authors are at the University of Washington and AI2. Verified 2026-08-13 via the Rowan/hellaswag dataset card on Hugging Face.

Openness

5 high confidence
5.0
data
open(59,950 examples, downloadable on HF)
license
MIT(permissive, redistributable)
datasheet
yes(HF dataset card + ACL 2019 paper)
access
public(train/val/test all released, not held-out)

Fully open under MIT with a dataset card - the canonical open-data benchmark profile. The repository carries no license tag, so the MIT is established by the body of the card rather than by a tag.

  • https://huggingface.co/datasets/Rowan/hellaswag recorded 2026-08-13

    the card's Licensing Information section reads "MIT" with a link to github.com/rowanz/hellaswag/blob/master/LICENSE; the embedded repo state reads `"gated":false` and the tags array carries no license tag; train/validation/test all public; 307,689 downloads in the trailing 30 days

Adoption

4 high confidence
4.0

309,927 Hugging Face downloads in the trailing 30 days for Rowan/hellaswag, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. HellaSwag is a commonsense benchmark cited across LLM release reports and on the Open LLM Leaderboard.

Capability

not assessed

Capability is not a meaningful axis for an eval dataset, so it is left unscored.

Verified 2026-08-12