AI Potluck
Model components / Benchmark / eval datasets

GSM8K

OpenAI

Grade School Math 8K: a benchmark of grade-school-level math word problems used to test chain-of-thought style reasoning and mathematical problem solving. It is a staple benchmark for evaluating small and frontier models on concise, single-answer math tasks.

GSM8K (openai/gsm8k), 8.5K grade-school math word problems (7,473 train / 1,319 test), ~5.9 MB, 'main' + 'socratic' configs. From 'Training Verifiers to Solve Math Word Problems' (arXiv:2110.14168). Verified live on HF June 2026 (~948k downloads last month).

Openness

5 high confidence
5.0
license
MIT(OSI-style permissive, fully redistributable)
redistributability
full(downloadable parquet, ungated)
datasheet
yes(dataset card)
size
8.5K problems
splits
train/test public

Fully open: MIT, ungated, downloadable, with a dataset card. Clean open-data case.

Adoption

4 high confidence
4.0

The standard grade-school math-reasoning benchmark, cited near-universally in LLM release reports for years. ~948k HF downloads/month on the canonical config (the highest in this batch) puts annualized usage comfortably into the 1-10M band.

Capability

not assessed

Dataset, not 'capable'. Now largely saturated by frontier models but per recipe capability is n/a.

Unchanged since 2026-06-24 (last edited, not re-checked)