AI Potluck
Model components / Training & synthetic datasets

Dolci

Allen Institute for AI

Ai2's post-training data suite for OLMo 3, covering supervised fine-tuning, DPO preference and RLVR sets for both the Instruct and Think variants. It is the successor to the Tulu mixtures, folding in OpenThoughts3, FLAN v2, OpenAssistant, WildChat and new Ai2 synthetic data.

A family record covering the Dolci-Instruct and Dolci-Think SFT, DPO and RL sets. Verified 2026-08-13 via the allenai/Dolci-Instruct-SFT dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, per-stage dataset cards.

Adoption

2 high confidence
2.0

3,807 downloads in the trailing 30 days for allenai/Dolci-Instruct-SFT, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

5 high confidence
5.0

The current fully-open frontier post-training suite; OLMo 3-Think 32B leads the fully-open thinking-model class in documented benchmarks. The Dolci card presents the suite as the successor to the Tulu mixtures.

Verified 2026-08-13