OpenWebMath
OpenWebMath projectCorpus of 14.7 billion tokens of mathematical web text extracted from Common Crawl with LaTeX preserved, released by researchers at the University of Toronto and collaborators. It appears in Proof-Pile-2 behind Llemma, and in OLMoE, Nemotron 3 and StarCoder2's math supplements.
Verified 2026-08-13 via the open-web-math/open-web-math dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY 1.0, ungated, full dataset card.
- https://huggingface.co/datasets/open-web-math/open-web-math recorded 2026-08-13
ODC-BY 1.0 license, ungated, dataset card
Adoption
3 high confidence33,285 downloads in the trailing 30 days for open-web-math/open-web-math, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/open-web-math/open-web-math recorded 2026-08-13
33,285 downloads in the trailing 30 days for open-web-math/open-web-math
Capability
3 high confidenceProven math corpus (Llemma/Proof-Pile-2, OLMoE, StarCoder2 math supplement) but Oct-2023 vintage and now the baseline that FineMath, MegaMath, and Nemotron-CC-Math beat in ablations.
- https://arxiv.org/html/2508.15096v1 recorded 2026-08-13
Nemotron-CC-Math surpasses all prior open math datasets including OpenWebMath
- https://huggingface.co/datasets/HuggingFaceTB/finemath recorded 2026-08-13
FineMath-3+ outperforms OpenWebMath on GSM8k and MATH
Verified 2026-08-13