AI Potluck
Model components / Benchmark / eval datasets

GPQA

Unknown

GPQA (Graduate-Level Google-Proof Q&A), 448 expert-written multiple-choice questions in biology/physics/chemistry; the 198-question 'GPQA Diamond' high-quality subset is the headline benchmark cited in most 2025-2026 frontier model reports. Paper arXiv:2311.12022. HF dataset Idavidrein/gpqa verified live June 2026.

GPQA (Graduate-Level Google-Proof Q&A), 448 expert-written multiple-choice questions in biology/physics/chemistry; the 198-question 'GPQA Diamond' high-quality subset is the headline benchmark cited in most 2025-2026 frontier model reports. Paper arXiv:2311.12022. HF dataset Idavidrein/gpqa verified live June 2026.

Openness

3 high confidence
3.0
license
CC-BY-4.0(open,redistributable+attribution)
access
gated(HF login + click-through agreement NOT to publish examples in plaintext, to limit training-corpus contamination)
datasheet
present(paper+card)
splits
public(no hidden held-out test)

Open CC-BY-4.0 license and fully redistributable, but access is login-walled with an anti-contamination click-through agreement, so the headline class is gated rather than open. Scored 3 with every other gated corpus on the map rather than 4 - a gate caps a dataset at 3 whatever its license says, because a corpus you need a login for is not more open than one you do not. The anti-contamination purpose is a fair argument for treating this gate as gentler than a restrictive one, but only this product records why its gate exists, and a rung for that would turn on evidence one product carries. Corrected 2026-08-01 from 4 when the dataset ladder made the pair checkable; no source was re-read.

Adoption

4 high confidence
4.0

~129,881 HF downloads in the trailing month (primary, from the dataset page); GPQA Diamond is a de-facto-standard reasoning benchmark cited across frontier model release reports (appears as a capability basis throughout the v3 flagship set), which is the top adoption signal for a benchmark.

Capability

not assessed

Per recipe, a benchmark dataset is not 'capable'; openness and adoption carry this category. No capability axis assessed.

Unchanged since 2026-08-01 (last edited, not re-checked)