AI Potluck
Model components / Training & synthetic datasets

Dolma

Allen Institute for AI

Open pretraining corpus of roughly 3 trillion tokens drawn from Common Crawl web text, academic publications, code, books and encyclopedic material. It was created by the Allen Institute for AI and used to train the OLMo family, and the release includes a toolkit for replicating and curating the data.

Superseded by Dolma 3 for OLMo 3. Verified 2026-08-13 via the allenai/dolma dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
ungated
yes
dataset_card
yes
code
apache-2.0-toolkit

Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.

Adoption

2 high confidence
2.0

2,724 downloads in the trailing 30 days for allenai/dolma, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

3 high confidence
3.0

Proven training corpus for OLMo 1/2 and OLMoE, but loses to FineWeb in the canonical FineWeb ablation and is superseded by Dolma 3.

Verified 2026-08-13