Tulu 3 SFT Mixture
Allen Institute for AIThe Tulu 3 SFT Mixture is a supervised fine-tuning dataset of roughly 939k samples combining 18 data sources, created by the Allen Institute for AI (Ai2) for its Tulu 3 post-training recipe. It covers conversation, instruction-following, math, code, safety, and multilingual data across 70+ languages. The mixture is ODC-BY-1.0 licensed while incorporating subsets under varying licenses. It underpins the open Tulu 3 model family.
Verified live 2026-06-22 via primary sources. ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.
Openness
5 high confidence- license
- odc-by(mixed subset licenses incl CC-BY-NC)
ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.
- https://huggingface.co/datasets/allenai/tulu-3-sft-mixture recorded 2026-06-22
ODC-BY-1.0 license, ~939k rows, dataset card
Adoption
3 high confidence18,317 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/allenai/tulu-3-sft-mixture recorded 2026-07-04
downloads field = 18,317 (30-day)
Capability
4 high confidenceCurrent multi-org SFT default (Marin, SmolLM3, OLMo 2) with SOTA open-recipe claims, but a reused component rather than the head-to-head leader; Dolci now leads the fully-open frontier.
- https://arxiv.org/abs/2411.15124 recorded 2026-07-04
The Tulu 3 recipe surpasses instruct versions of Llama 3.1, Qwen 2.5, and GPT-4o-mini
- https://huggingface.co/marin-community/marin-8b-instruct recorded 2026-07-04
Marin 8B Instruct lists tulu-3-sft-mixture among its SFT sources
Unchanged since 2026-07-04 (last edited, not re-checked)