AI Potluck
Model components / Training & synthetic datasets

Dolma

Allen Institute for AI

Dolma is an open pretraining corpus of roughly 3 trillion tokens drawn from Common Crawl web text, academic publications, code, books, and encyclopedic material. It was created by the Allen Institute for AI (AI2) and used to train the OLMo family of models. The release includes the corpus plus an open source toolkit for replicating and curating the data.

Verified live 2026-06-22 via primary sources. Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.

Openness

5 high confidence
5.0
license
ODC-BY
ungated
yes
dataset_card
yes
code
apache-2.0-toolkit

Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.

Adoption

2 high confidence
2.0

4,015 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Proven training corpus for OLMo 1/2 and OLMoE, but loses to FineWeb in the canonical FineWeb ablation and is superseded by Dolma 3.

Unchanged since 2026-07-04 (last edited, not re-checked)