FLAN Collection (v2)
Google ResearchThe FLAN Collection (v2) is Google Research's instruction-tuning compilation of 1,800+ tasks (P3, Super-NaturalInstructions, chain-of-thought data) that trained Flan-T5 and Flan-PaLM. It remains a canonical instruction-data building block, folded into Tulu 3, OLMo 3's Dolci, and Marin's Dolmino mixes.
Openness
4 medium confidence- access
- ungated(conversion)
- license
- Apache-2.0(recipe)+per-task
- dataset_card
- partial
Released as an Apache-2.0 recipe/repo; the widely used HF conversion (ai2-adapt-dev) is ungated. Underlying task datasets keep their own licenses.
- https://github.com/google-research/FLAN recorded 2026-07-04
Apache-2.0 FLAN v2 recipe
- https://huggingface.co/datasets/ai2-adapt-dev/flan_v2_converted recorded 2026-07-04
ungated HF conversion used by Tulu/OLMo mixes
Adoption
1 high confidence924 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/ai2-adapt-dev/flan_v2_converted recorded 2026-07-04
downloads field = 924 (30-day)
Capability
3 high confidenceStrong historical comparative ablations (Flan 2022 beats P3/T0/SNI on the same T5-XL) and still a component in Tulu 3, Dolci, and Marin's mixes, but dated as a standalone chat-era recipe.
- https://arxiv.org/abs/2301.13688 recorded 2026-07-04
The Flan Collection: Flan 2022 leads P3++/Super-NatInst on MMLU/BBH (same T5-XL)
- https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT recorded 2026-07-04
OLMo 3 Dolci still consumes FLAN v2 (89,981 prompts)
Unchanged since 2026-07-04 (last edited, not re-checked)