Tulu 3 SFT Mixture
Allen Institute for AISupervised fine-tuning dataset of roughly 939,000 samples combining 18 data sources, created by the Allen Institute for AI for its Tulu 3 post-training recipe. It covers conversation, instruction-following, math, code, safety and multilingual data across more than 70 languages, and underpins the Tulu 3 model family.
Some constituent subsets carry terms narrower than the mixture's own, which the card documents. Verified 2026-08-13 via the allenai/tulu-3-sft-mixture dataset card on Hugging Face.
Openness
5 high confidence- license
- odc-by(mixed subset licenses incl CC-BY-NC)
- access
- ungated
- dataset_card
- present
ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.
- https://huggingface.co/datasets/allenai/tulu-3-sft-mixture recorded 2026-08-13
ODC-BY-1.0 license, ~939k rows, dataset card
Adoption
3 high confidence36,267 downloads in the trailing 30 days for allenai/tulu-3-sft-mixture, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/tulu-3-sft-mixture recorded 2026-08-13
36,267 downloads in the trailing 30 days for allenai/tulu-3-sft-mixture
Capability
4 high confidenceCurrent multi-org SFT default (Marin, SmolLM3, OLMo 2) with SOTA open-recipe claims, but a reused component rather than the head-to-head leader; Dolci now leads the fully-open frontier. The Tulu 3 listing claims results surpassing the instruct versions of Llama 3.1, Qwen 2.5, Mistral and GPT-4o-mini.
- https://arxiv.org/abs/2411.15124 recorded 2026-08-13
The Tulu 3 recipe surpasses instruct versions of Llama 3.1, Qwen 2.5, and GPT-4o-mini
- https://huggingface.co/marin-community/marin-8b-instruct recorded 2026-08-13
Marin 8B Instruct lists tulu-3-sft-mixture among its SFT sources
Verified 2026-08-13