AI Potluck
Model components / Training & synthetic datasets

Dolci

Allen Institute for AI

Dolci is Ai2's post-training data suite for OLMo 3: SFT, DPO preference, and RLVR sets for both Instruct and Think variants, released under ODC-BY. It is the successor to the Tulu mixtures, folding in OpenThoughts3, FLAN v2, OpenAssistant, WildChat, and new Ai2 synthetic data.

Family bucket: Dolci-Instruct and Dolci-Think SFT/DPO/RL sets.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, per-stage dataset cards.

Adoption

3 high confidence
3.0

12,526 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

5 high confidence
5.0

The current fully-open frontier post-training suite; OLMo 3-Think 32B leads the fully-open thinking-model class in documented benchmarks.

Unchanged since 2026-07-04 (last edited, not re-checked)