Dolma 3
Allen Institute for AIAi2's pretraining data release behind OLMo 3: a pool of about 9 trillion tokens drawn from Common Crawl, olmOCR-processed scientific PDFs, Stack-Edu code and FineMath, alongside the quality-upsampled mixes used in training - Dolma 3 Mix, Dolmino mid-training and Longmino long-context. The pool, every mix and the curation tooling are released together.
A family record covering the pool and the training mixes; successor to the map's Dolma 1.x entry. Verified 2026-08-13 via the allenai/dolma3_pool dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, fully documented pool and training mixes.
- https://huggingface.co/datasets/allenai/dolma3_pool recorded 2026-08-13
ODC-BY license, ungated, dataset card
Adoption
3 high confidence36,467 downloads in the trailing 30 days for allenai/dolma3_pool, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/dolma3_pool recorded 2026-08-13
36,467 downloads in the trailing 30 days for allenai/dolma3_pool
Capability
5 high confidenceThe current (2025) corpus behind OLMo 3, the strongest fully-open base and thinking models; uses documented quality-aware upsampling and data-mix ablations.
- https://allenai.org/blog/olmo3 recorded 2026-08-13
Dolma 3 trains all OLMo 3 variants; OLMo 3-Base 32B is the strongest fully-open base model
Verified 2026-08-13