AI Potluck
Model components / Benchmark / eval datasets

MATH

Hendrycks et al. (UC Berkeley)

Benchmark of 12,500 competition mathematics problems spanning algebra, counting and probability, geometry, number theory, prealgebra and precalculus, split 7,500 train and 5,000 test, each with a step-by-step solution in LaTeX and a difficulty level from 1 to 5. It measures reasoning depth beyond grade-school word problems.

The canonical Hugging Face repository is access-disabled under a DMCA takedown, so the data is no longer redistributable from the registry and the trailing-30-day download count is zero. Verified 2026-08-13 via the hendrycks/competition_math dataset page.

Openness

1 medium confidence
1.0
license
MIT(declared)
access
closed(canonical HF repo access-disabled under a DMCA takedown)
data
NOT downloadable from canonical HF repo (access disabled, DMCA takedown)
datasheet
present(card still documents schema)
redistributability
BLOCKED on the primary registry

Declared MIT, but the canonical Hugging Face dataset is access-disabled under a DMCA takedown over AMC/AIME problem rights, so the corpus is no longer distributed from its primary registry. Data that is not distributed at all counts as closed rather than gated, whatever the license says and however it came to be that way. The Hub API reports 0 downloads in the trailing 30 days against 1,034,712 all-time, which is what a disabled repository looks like. Mirrors and eval harnesses still embed the problems, so this is worth revisiting if the takedown is resolved.

  • https://huggingface.co/datasets/hendrycks/competition_math recorded 2026-08-13

    "Access to this dataset has been disabled" with "DMCA Takedown notice - see huggingface.co/datasets/hendrycks/competition_math/discussions/5"; the `license:mit` tag and the "Mathematics Aptitude Test of Heuristics" dataset card are still rendered; the embedded repo state reports 0 downloads in the trailing 30 days

Adoption

4 medium confidence
4.0

MATH is one of the most-cited math-reasoning standards, and this band rests on its standing in release reports rather than on a download count. The canonical Hugging Face repository is takedown-walled and reports 0 downloads in the trailing 30 days against 1,034,712 all-time, so no usage figure can be obtained from it. The standing is evidenced by OLMo 3's official model card, whose evaluation table carries MATH as a row - 65.1 through 91.6 across nine compared models - alongside AIME 2024/2025, GPQA, MMLU, IFEval, LiveCodeBench and thirteen others: a frontier open model reporting MATH as a headline metric in its release evaluation. The takedown page is cited beside it because it is what rules a measured download figure out.

  • https://huggingface.co/allenai/Olmo-3-7B-Instruct/raw/main/README.md recorded 2026-08-13

    evaluation table row "| **Math** | MATH | 65.1 | 79.6 | 87.3 | 82.3 | 91.6 | 71.0 | 30.1 | 21.9 | 67.3 |", reported beside AIME 2024, AIME 2025, OMEGA, BigBenchHard, GPQA, MMLU, IFEval, LiveCodeBench v3 and others - MATH carried as a headline metric in a frontier open model release

  • https://huggingface.co/datasets/hendrycks/competition_math recorded 2026-08-13

    the access-disabled page and its DMCA notice; carries the `license:mit` tag and the card, and reports 0 downloads in the trailing 30 days. This is what rules out a usage_volume band; it establishes nothing about release-report standing

Capability

not assessed

A dataset is not 'capable', so this axis is left unscored.

Verified 2026-08-13