AI Potluck
Model components / Benchmark / eval datasets

MATH

Hendrycks et al. (UC Berkeley)

A 12,500-problem benchmark of competition-style mathematics questions spanning algebra, counting, geometry, number theory, prealgebra, precalculus, and more. It is the standard hard math benchmark for measuring reasoning depth beyond GSM8K.

MATH (Hendrycks et al., NeurIPS 2021, arXiv 2103.03874): 12,500 competition math problems (7,500 train / 5,000 test) with step-by-step LaTeX solutions, Level 1-5, 7 subjects. Original HF dataset (hendrycks/competition_math) verified June 2026 but ACCESS DISABLED via DMCA takedown; no longer redistributable from the canonical registry; license tag is MIT.

Openness

2 medium confidence
2.0
data
NOT downloadable from canonical HF repo (access disabled, DMCA takedown)
license
MIT(declared)
datasheet
present(card still documents schema)
redistributability
BLOCKED on the primary registry

Declared MIT, but the canonical HF dataset is access-disabled under a DMCA takedown, so it is not currently redistributable from the primary registry; scored gated/blocked rather than open. Mirrors and eval harnesses still embed it, but the primary source is walled; revisit if the takedown is resolved.

Adoption

4 medium confidence
4.0

MATH is one of the most-cited math-reasoning standards; MATH / MATH-500 appears as a headline metric across frontier model release reports (the top adoption signal for a benchmark). Despite the canonical HF repo being takedown-walled, the benchmark itself remains pervasively used via mirrors and eval harnesses. Level 4 on release-report standard status; no clean current download count obtainable from the primary repo (disabled), hence reported_traction not usage_volume.

Capability

not assessed

Dataset, not capability-scored.

Unchanged since 2026-06-24 (last edited, not re-checked)