Dolma
Allen Institute for AIDolma is an open pretraining corpus of roughly 3 trillion tokens drawn from Common Crawl web text, academic publications, code, books, and encyclopedic material. It was created by the Allen Institute for AI (AI2) and used to train the OLMo family of models. The release includes the corpus plus an open source toolkit for replicating and curating the data.
Verified live 2026-06-22 via primary sources. Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.
Openness
5 high confidence- license
- ODC-BY
- ungated
- yes
- dataset_card
- yes
- code
- apache-2.0-toolkit
Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.
- https://huggingface.co/datasets/allenai/dolma recorded 2026-06-22
ODC-BY license, ungated access, dataset card, ~3T tokens
Adoption
2 high confidence4,015 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/allenai/dolma recorded 2026-07-04
downloads field = 4,015 (30-day)
Capability
3 high confidenceProven training corpus for OLMo 1/2 and OLMoE, but loses to FineWeb in the canonical FineWeb ablation and is superseded by Dolma 3.
- https://arxiv.org/abs/2406.17557 recorded 2026-07-04
FineWeb outperforms Dolma v1.6 and v1.7 in the 1.82B/350B ablation
- https://allenai.org/blog/olmo3 recorded 2026-07-04
Dolma superseded by Dolma 3 for OLMo 3
Unchanged since 2026-07-04 (last edited, not re-checked)