Dolma 3
Allen Institute for AIDolma 3 is Ai2's ODC-BY pretraining data release behind OLMo 3: a ~9T-token pool from Common Crawl, olmOCR-processed scientific PDFs, Stack-Edu code, and FineMath, plus the exact quality-upsampled mixes (Dolma 3 Mix, Dolmino mid-training, Longmino long-context) used in training. The full pool, every mix, and the curation tooling are public.
Family bucket: dolma3_pool plus the Mix/Dolmino/Longmino training mixes. Successor to the map's dolma entry (Dolma 1.x).
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, fully documented pool and training mixes.
- https://huggingface.co/datasets/allenai/dolma3_pool recorded 2026-07-04
ODC-BY license, ungated, dataset card
Adoption
4 high confidence158,294 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/allenai/dolma3_mix-6T-1025-7B recorded 2026-07-04
downloads field = 158,294 (30-day)
Capability
5 high confidenceThe current (2025) corpus behind OLMo 3, the strongest fully-open base and thinking models; uses documented quality-aware upsampling and data-mix ablations.
- https://allenai.org/blog/olmo3 recorded 2026-07-04
Dolma 3 trains all OLMo 3 variants; OLMo 3-Base 32B is the strongest fully-open base model
Unchanged since 2026-07-04 (last edited, not re-checked)