Dolma
Allen Institute for AIOpen pretraining corpus of roughly 3 trillion tokens drawn from Common Crawl web text, academic publications, code, books and encyclopedic material. It was created by the Allen Institute for AI and used to train the OLMo family, and the release includes a toolkit for replicating and curating the data.
Superseded by Dolma 3 for OLMo 3. Verified 2026-08-13 via the allenai/dolma dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- ungated
- yes
- dataset_card
- yes
- code
- apache-2.0-toolkit
Permissively licensed (ODC-BY), ungated download with a full dataset card and replication toolkit.
- https://huggingface.co/datasets/allenai/dolma recorded 2026-08-13
ODC-BY license, ungated access, dataset card, ~3T tokens
Adoption
2 high confidence2,724 downloads in the trailing 30 days for allenai/dolma, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/dolma recorded 2026-08-13
2,724 downloads in the trailing 30 days for allenai/dolma
Capability
3 high confidenceProven training corpus for OLMo 1/2 and OLMoE, but loses to FineWeb in the canonical FineWeb ablation and is superseded by Dolma 3.
- https://arxiv.org/abs/2406.17557 recorded 2026-08-13
Abstract - FineWeb produces better-performing LLMs than other open pretraining datasets. The Dolma v1.6/v1.7 ablation rows are in the paper body, not on this page
- https://allenai.org/blog/olmo3 recorded 2026-08-13
Dolma superseded by Dolma 3 for OLMo 3
Verified 2026-08-13