Tulu 3
Ai2Ai2's open post-training model family (Llama-3.1-Tulu-3 8B / 70B / 405B), the reference fully-open post-training recipe: open SFT + preference data, open training code (open-instruct), and the RLVR (Reinforcement Learning with Verifiable Rewards) method, applied on top of Meta's Llama 3.1 base. Note: the recipe/data/code are Apache-2.0 and fully reproducible, but the released weights inherit Meta's Llama 3.1 Community License (non-OSI), so the artifact is open-recipe over restricted weights. Verified live on HF June 2026.
Weights under Llama 3.1 Community License (non-OSI); Ai2's post-training data, code, recipe, and eval framework (open-instruct, OLMES) are Apache-2.0. Classed as restricted on the weights license, consistent with how other Llama-Community models are scored.
Openness
3 high confidence- weights
- open(downloadable)
- license
- Llama-3.1-Community(non-OSI
- post-training-data
- open(Apache-2.0/ODC-BY)
- code
- open(open-instruct Apache-2.0)
- recipe
- open(SFT->DPO->RLVR)
- evals
- open(OLMES)
The released weights inherit Meta's Llama 3.1 Community License (non-OSI, 700M-MAU cap), so the license governs what a user downloading THIS checkpoint is bound by, not Ai2's post- training layer. That layer is genuinely open - data, code, RLVR recipe and eval framework, all Apache-2.0 - and it is credited where it lives: tulu-3-sft-mixture carries 5/open in training_synthetic_datasets, and the same recipe produces olmo-3-instruct at 5/open_source on an open base. Scoring the method on the checkpoint would score two things at once. Was 3, corrected to 2 on 2026-07-29, and now 3 again on 2026-08-01 - but not a reversal to the original reasoning. The 2026-07-29 correction was right at the time: the 3 it removed was 3/restricted, a pair no rule in the category ladder could produce. The universal license scale now caps a bounded commercial license at 3/open_weights, which IS producible, so the score returns while the objection that removed it stays answered.
- https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B recorded 2026-06-25
License: Llama 3.1 Community License Agreement (non-OSI); base meta-llama/Llama-3.1-8B
- https://allenai.org/blog/tulu-3 recorded 2026-06-25
fully-transparent post-training release -- open data + data mixes, training code/recipe, eval framework, and weights
- https://github.com/allenai/open-instruct recorded 2026-06-25
Apache-2.0; Tulu 3 SFT/DPO/RLVR post-training framework + decontamination tools
Adoption
2 high confidenceThe released weights inherit Meta's Llama 3.1 Community License (non-OSI, 700M-MAU cap), so the license governs what a user downloading THIS checkpoint is bound by, not Ai2's post- training layer. That layer is genuinely open - data, code, RLVR recipe and eval framework, all Apache-2.0 - and it is credited where it lives: tulu-3-sft-mixture carries 5/open in training_synthetic_datasets, and the same recipe produces olmo-3-instruct at 5/open_source on an open base. Scoring the method on the checkpoint would score two things at once. Was 3, corrected to 2 on 2026-07-29, and now 3 again on 2026-08-01 - but not a reversal to the original reasoning. The 2026-07-29 correction was right at the time: the 3 it removed was 3/restricted, a pair no rule in the category ladder could produce. The universal license scale now caps a bounded commercial license at 3/open_weights, which IS producible, so the score returns while the objection that removed it stays answered.
- https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B recorded 2026-06-25
~3,109 downloads last month; 179 likes
- https://huggingface.co/allenai/Llama-3.1-Tulu-3-405B recorded 2026-06-25
~832 downloads last month; 112 likes
Capability
3 high confidenceThe released weights inherit Meta's Llama 3.1 Community License (non-OSI, 700M-MAU cap), so the license governs what a user downloading THIS checkpoint is bound by, not Ai2's post- training layer. That layer is genuinely open - data, code, RLVR recipe and eval framework, all Apache-2.0 - and it is credited where it lives: tulu-3-sft-mixture carries 5/open in training_synthetic_datasets, and the same recipe produces olmo-3-instruct at 5/open_source on an open base. Scoring the method on the checkpoint would score two things at once. Was 3, corrected to 2 on 2026-07-29, and now 3 again on 2026-08-01 - but not a reversal to the original reasoning. The 2026-07-29 correction was right at the time: the 3 it removed was 3/restricted, a pair no rule in the category ladder could produce. The universal license scale now caps a bounded commercial license at 3/open_weights, which IS producible, so the score returns while the objection that removed it stays answered.
- https://huggingface.co/allenai/Llama-3.1-Tulu-3-405B recorded 2026-06-25
MMLU 87.0, GSM8K 95.5, IFEval 86.0, HumanEval 95.9, MATH 67.3; vs GPT-4o comparison
- https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B recorded 2026-06-25
MMLU 68.2, GSM8K 87.6, IFEval 82.4
Unchanged since 2026-08-07 (last edited, not re-checked)