AI Potluck
Model components / Training & synthetic datasets

Dolma 3

Allen Institute for AI

Dolma 3 is Ai2's ODC-BY pretraining data release behind OLMo 3: a ~9T-token pool from Common Crawl, olmOCR-processed scientific PDFs, Stack-Edu code, and FineMath, plus the exact quality-upsampled mixes (Dolma 3 Mix, Dolmino mid-training, Longmino long-context) used in training. The full pool, every mix, and the curation tooling are public.

Family bucket: dolma3_pool plus the Mix/Dolmino/Longmino training mixes. Successor to the map's dolma entry (Dolma 1.x).

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, fully documented pool and training mixes.

Adoption

4 high confidence
4.0

158,294 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

5 high confidence
5.0

The current (2025) corpus behind OLMo 3, the strongest fully-open base and thinking models; uses documented quality-aware upsampling and data-mix ablations.

Unchanged since 2026-07-04 (last edited, not re-checked)