GPQA
UnknownGPQA (Graduate-Level Google-Proof Q&A), 448 expert-written multiple-choice questions in biology/physics/chemistry; the 198-question 'GPQA Diamond' high-quality subset is the headline benchmark cited in most 2025-2026 frontier model reports. Paper arXiv:2311.12022. HF dataset Idavidrein/gpqa verified live June 2026.
GPQA (Graduate-Level Google-Proof Q&A), 448 expert-written multiple-choice questions in biology/physics/chemistry; the 198-question 'GPQA Diamond' high-quality subset is the headline benchmark cited in most 2025-2026 frontier model reports. Paper arXiv:2311.12022. HF dataset Idavidrein/gpqa verified live June 2026.
Openness
3 high confidence- license
- CC-BY-4.0(open,redistributable+attribution)
- access
- gated(HF login + click-through agreement NOT to publish examples in plaintext, to limit training-corpus contamination)
- datasheet
- present(paper+card)
- splits
- public(no hidden held-out test)
Open CC-BY-4.0 license and fully redistributable, but access is login-walled with an anti-contamination click-through agreement, so the headline class is gated rather than open. Scored 3 with every other gated corpus on the map rather than 4 - a gate caps a dataset at 3 whatever its license says, because a corpus you need a login for is not more open than one you do not. The anti-contamination purpose is a fair argument for treating this gate as gentler than a restrictive one, but only this product records why its gate exists, and a rung for that would turn on evidence one product carries. Corrected 2026-08-01 from 4 when the dataset ladder made the pair checkable; no source was re-read.
- https://huggingface.co/datasets/Idavidrein/gpqa recorded 2026-06-04
CC-BY-4.0 license, gated access with anti-leakage agreement, 448 questions, dataset card present
- https://arxiv.org/abs/2311.12022 recorded 2026-06-04
GPQA paper documenting 448 expert questions + Diamond subset
Adoption
4 high confidence~129,881 HF downloads in the trailing month (primary, from the dataset page); GPQA Diamond is a de-facto-standard reasoning benchmark cited across frontier model release reports (appears as a capability basis throughout the v3 flagship set), which is the top adoption signal for a benchmark.
- https://huggingface.co/datasets/Idavidrein/gpqa recorded 2026-06-04
129,881 downloads last month
Capability
not assessedPer recipe, a benchmark dataset is not 'capable'; openness and adoption carry this category. No capability axis assessed.
Unchanged since 2026-08-01 (last edited, not re-checked)