Dolci
Allen Institute for AIDolci is Ai2's post-training data suite for OLMo 3: SFT, DPO preference, and RLVR sets for both Instruct and Think variants, released under ODC-BY. It is the successor to the Tulu mixtures, folding in OpenThoughts3, FLAN v2, OpenAssistant, WildChat, and new Ai2 synthetic data.
Family bucket: Dolci-Instruct and Dolci-Think SFT/DPO/RL sets.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, per-stage dataset cards.
- https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT recorded 2026-07-04
ODC-BY license, ungated, dataset card
Adoption
3 high confidence12,526 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/allenai/Dolci-Instruct-SFT recorded 2026-07-04
downloads field = 12,526 (30-day)
Capability
5 high confidenceThe current fully-open frontier post-training suite; OLMo 3-Think 32B leads the fully-open thinking-model class in documented benchmarks.
- https://allenai.org/blog/olmo3 recorded 2026-07-04
OLMo 3-Think 32B (trained with Dolci) is the strongest fully-open thinking model: MATH 96.1, HumanEval+ 91.4, IFEval 89.0
- https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT recorded 2026-07-04
Dolci is Ai2's OLMo 3 post-training suite, successor to the Tulu mixtures
Unchanged since 2026-07-04 (last edited, not re-checked)