AI Potluck
Model components / Training & synthetic datasets

OpenWebMath

OpenWebMath project

Corpus of 14.7 billion tokens of mathematical web text extracted from Common Crawl with LaTeX preserved, released by researchers at the University of Toronto and collaborators. It appears in Proof-Pile-2 behind Llemma, and in OLMoE, Nemotron 3 and StarCoder2's math supplements.

Verified 2026-08-13 via the open-web-math/open-web-math dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY 1.0, ungated, full dataset card.

Adoption

3 high confidence
3.0

33,285 downloads in the trailing 30 days for open-web-math/open-web-math, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Proven math corpus (Llemma/Proof-Pile-2, OLMoE, StarCoder2 math supplement) but Oct-2023 vintage and now the baseline that FineMath, MegaMath, and Nemotron-CC-Math beat in ablations.

Verified 2026-08-13