HellaSwag
Allen Institute for AICommonsense natural-language-inference and sentence-completion benchmark from 2019, holding 59,950 four-way multiple-choice examples over contexts derived from ActivityNet. It tests whether a model picks the plausible continuation of an everyday situation, and is reported in most language-model releases.
The registry attributes the dataset to the Allen Institute; the primary authors are at the University of Washington and AI2. Verified 2026-08-13 via the Rowan/hellaswag dataset card on Hugging Face.
Openness
5 high confidence- data
- open(59,950 examples, downloadable on HF)
- license
- MIT(permissive, redistributable)
- datasheet
- yes(HF dataset card + ACL 2019 paper)
- access
- public(train/val/test all released, not held-out)
Fully open under MIT with a dataset card - the canonical open-data benchmark profile. The repository carries no license tag, so the MIT is established by the body of the card rather than by a tag.
- https://huggingface.co/datasets/Rowan/hellaswag recorded 2026-08-13
the card's Licensing Information section reads "MIT" with a link to github.com/rowanz/hellaswag/blob/master/LICENSE; the embedded repo state reads `"gated":false` and the tags array carries no license tag; train/validation/test all public; 307,689 downloads in the trailing 30 days
Adoption
4 high confidence309,927 Hugging Face downloads in the trailing 30 days for Rowan/hellaswag, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. HellaSwag is a commonsense benchmark cited across LLM release reports and on the Open LLM Leaderboard.
- https://huggingface.co/api/datasets/Rowan/hellaswag recorded 2026-08-12
309,927 downloads in the trailing 30 days for Rowan/hellaswag
Capability
not assessedCapability is not a meaningful axis for an eval dataset, so it is left unscored.
- https://huggingface.co/datasets/Rowan/hellaswag recorded 2026-08-13
a static 4-way multiple-choice completion corpus; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12