AI Potluck
Model components / Training & synthetic datasets

Tulu 3 SFT Mixture

Allen Institute for AI

The Tulu 3 SFT Mixture is a supervised fine-tuning dataset of roughly 939k samples combining 18 data sources, created by the Allen Institute for AI (Ai2) for its Tulu 3 post-training recipe. It covers conversation, instruction-following, math, code, safety, and multilingual data across 70+ languages. The mixture is ODC-BY-1.0 licensed while incorporating subsets under varying licenses. It underpins the open Tulu 3 model family.

Verified live 2026-06-22 via primary sources. ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.

Openness

5 high confidence
5.0
license
odc-by(mixed subset licenses incl CC-BY-NC)

ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.

Adoption

3 high confidence
3.0

18,317 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

Current multi-org SFT default (Marin, SmolLM3, OLMo 2) with SOTA open-recipe claims, but a reused component rather than the head-to-head leader; Dolci now leads the fully-open frontier.

Unchanged since 2026-07-04 (last edited, not re-checked)