GSM8K
OpenAIGrade School Math 8K: 8,500 grade-school math word problems, 7,473 for training and 1,319 for test, used to evaluate chain-of-thought reasoning on concise single-answer tasks. It ships main and socratic configurations and comes from the paper on training verifiers to solve math word problems.
Verified 2026-08-13 via the openai/gsm8k dataset card on Hugging Face.
Openness
5 high confidence- license
- MIT(OSI-style permissive, fully redistributable)
- redistributability
- full(downloadable parquet, ungated)
- datasheet
- yes(dataset card)
- size
- 8.5K problems
- splits
- train/test public
Fully open: MIT, ungated, downloadable, with a dataset card. Clean open-data case.
- https://huggingface.co/datasets/openai/gsm8k recorded 2026-08-13
`license:mit` in the repo tags; the embedded repo state reads `"gated":false`; the dataset card renders with the main and socratic configs; 974,089 downloads in the trailing 30 days
Adoption
4 high confidence960,715 Hugging Face downloads in the trailing 30 days for openai/gsm8k, which puts it in the 100K-1M band, level 4 on the dataset adoption scale. That scale tops out at >1M, an order of magnitude below the ones used for software and models, because no dataset in this corpus has ever passed 10M downloads. GSM8K is the standard grade-school math-reasoning benchmark, cited near-universally in LLM release reports.
- https://huggingface.co/api/datasets/openai/gsm8k recorded 2026-08-12
960,715 downloads in the trailing 30 days for openai/gsm8k
Capability
not assessedA dataset is not 'capable', so this axis is left unscored. GSM8K is now largely saturated by frontier models, but that is a quality caveat rather than a capability score.
- https://huggingface.co/datasets/openai/gsm8k recorded 2026-08-13
a static corpus of grade-school word problems with question and answer columns; no performance, throughput or feature claim on the page for the capability axis to read
Verified 2026-08-12