GPQA
UnknownGraduate-level Google-proof question answering: 448 expert-written multiple-choice questions across biology, physics and chemistry, written so that they resist quick web lookup. A 198-question Diamond subset carries the highest-agreement items and is the figure most frontier model reports cite.
Verified 2026-08-13 via the Idavidrein/gpqa dataset card on Hugging Face and the GPQA paper abstract.
Openness
3 high confidence- license
- CC-BY-4.0(open,redistributable+attribution)
- access
- gated(HF login + click-through agreement NOT to publish examples in plaintext, to limit training-corpus contamination)
- datasheet
- present(paper+card)
- splits
- public(no hidden held-out test)
The license is CC-BY-4.0 and the corpus is fully redistributable, but access is login-walled behind an anti-contamination click-through agreement, so the headline class is gated rather than open. A gate caps a dataset at 3 whatever its license says, because a corpus you need a login for is not more open than one you do not, and that puts GPQA level with every other gated corpus on the map. The anti-contamination purpose is a fair argument for treating this gate as gentler than a restrictive one, but GPQA is the only product here that records why its gate exists, so a rung for that would turn on evidence a single product happens to carry.
- https://huggingface.co/datasets/Idavidrein/gpqa recorded 2026-08-13
`license:cc-by-4.0` in the repo tags; the embedded repo state reads `"gated":"auto"` and the page still says "You need to agree to share your contact information to access this dataset"; the dataset card still renders
- https://arxiv.org/abs/2311.12022 recorded 2026-08-13
the abstract still reads "a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry"
Adoption
4 high confidence110,577 Hugging Face downloads in the trailing 30 days across the declared artifacts, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. GPQA Diamond is a de-facto standard reasoning benchmark cited across frontier model release reports, and it appears as the basis for capability scores throughout the flagship models on this map - the strongest adoption signal a benchmark can have.
- https://huggingface.co/api/datasets/Idavidrein/gpqa recorded 2026-08-13
110,577 downloads in the trailing 30 days for Idavidrein/gpqa
Capability
not assessedA benchmark dataset is not 'capable', so this axis is left unscored; openness and adoption carry this category.
- https://huggingface.co/datasets/Idavidrein/gpqa recorded 2026-08-13
a static expert-written multiple-choice corpus; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-13