AI Potluck
Model components / Benchmark / eval datasets

GPQA

Unknown

Graduate-level Google-proof question answering: 448 expert-written multiple-choice questions across biology, physics and chemistry, written so that they resist quick web lookup. A 198-question Diamond subset carries the highest-agreement items and is the figure most frontier model reports cite.

Verified 2026-08-13 via the Idavidrein/gpqa dataset card on Hugging Face and the GPQA paper abstract.

Openness

3 high confidence
3.0
license
CC-BY-4.0(open,redistributable+attribution)
access
gated(HF login + click-through agreement NOT to publish examples in plaintext, to limit training-corpus contamination)
datasheet
present(paper+card)
splits
public(no hidden held-out test)

The license is CC-BY-4.0 and the corpus is fully redistributable, but access is login-walled behind an anti-contamination click-through agreement, so the headline class is gated rather than open. A gate caps a dataset at 3 whatever its license says, because a corpus you need a login for is not more open than one you do not, and that puts GPQA level with every other gated corpus on the map. The anti-contamination purpose is a fair argument for treating this gate as gentler than a restrictive one, but GPQA is the only product here that records why its gate exists, so a rung for that would turn on evidence a single product happens to carry.

  • https://huggingface.co/datasets/Idavidrein/gpqa recorded 2026-08-13

    `license:cc-by-4.0` in the repo tags; the embedded repo state reads `"gated":"auto"` and the page still says "You need to agree to share your contact information to access this dataset"; the dataset card still renders

  • https://arxiv.org/abs/2311.12022 recorded 2026-08-13

    the abstract still reads "a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry"

Adoption

4 high confidence
4.0

110,577 Hugging Face downloads in the trailing 30 days across the declared artifacts, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. GPQA Diamond is a de-facto standard reasoning benchmark cited across frontier model release reports, and it appears as the basis for capability scores throughout the flagship models on this map - the strongest adoption signal a benchmark can have.

Capability

not assessed

A benchmark dataset is not 'capable', so this axis is left unscored; openness and adoption carry this category.

Verified 2026-08-13