AI Potluck
Model components / Training & synthetic datasets

Dolma 3

Allen Institute for AI

Ai2's pretraining data release behind OLMo 3: a pool of about 9 trillion tokens drawn from Common Crawl, olmOCR-processed scientific PDFs, Stack-Edu code and FineMath, alongside the quality-upsampled mixes used in training - Dolma 3 Mix, Dolmino mid-training and Longmino long-context. The pool, every mix and the curation tooling are released together.

A family record covering the pool and the training mixes; successor to the map's Dolma 1.x entry. Verified 2026-08-13 via the allenai/dolma3_pool dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, fully documented pool and training mixes.

Adoption

3 high confidence
3.0

36,467 downloads in the trailing 30 days for allenai/dolma3_pool, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

5 high confidence
5.0

The current (2025) corpus behind OLMo 3, the strongest fully-open base and thinking models; uses documented quality-aware upsampling and data-mix ablations.

Verified 2026-08-13