AI Potluck
Model components / Training & synthetic datasets

Tulu 3 SFT Mixture

Allen Institute for AI

Supervised fine-tuning dataset of roughly 939,000 samples combining 18 data sources, created by the Allen Institute for AI for its Tulu 3 post-training recipe. It covers conversation, instruction-following, math, code, safety and multilingual data across more than 70 languages, and underpins the Tulu 3 model family.

Some constituent subsets carry terms narrower than the mixture's own, which the card documents. Verified 2026-08-13 via the allenai/tulu-3-sft-mixture dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
odc-by(mixed subset licenses incl CC-BY-NC)
access
ungated
dataset_card
present

ODC-BY licensed and ungated with a full dataset card, though some subsets carry non-commercial licenses.

Adoption

3 high confidence
3.0

36,267 downloads in the trailing 30 days for allenai/tulu-3-sft-mixture, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

4 high confidence
4.0

Current multi-org SFT default (Marin, SmolLM3, OLMo 2) with SOTA open-recipe claims, but a reused component rather than the head-to-head leader; Dolci now leads the fully-open frontier. The Tulu 3 listing claims results surpassing the instruct versions of Llama 3.1, Qwen 2.5, Mistral and GPT-4o-mini.

Verified 2026-08-13