FineWeb-Edu
Hugging FaceFineWeb-Edu is an educational subset of FineWeb (~1.3 trillion tokens) created by Hugging Face's FineWeb team. It was produced by filtering FineWeb with an educational-quality classifier trained on Llama-3-70B annotations, retaining text scored as highly educational. It is intended to improve LLM performance on knowledge- and reasoning-heavy benchmarks relative to unfiltered web data.
Verified live 2026-06-22 via primary sources. Permissive ODC-BY license with ungated download and full dataset card.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- format
- parquet
Permissive ODC-BY license with ungated download and full dataset card.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu recorded 2026-06-22
Dataset card present, ODC-BY license, public ungated download
Adoption
4 high confidence384,910 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb-edu recorded 2026-07-04
downloads field = 384,910 (30-day)
Capability
5 high confidenceDominates FineWeb and all peers on MMLU/ARC, matching full FineWeb with ~10x fewer tokens; current default web corpus for SmolLM2/3 and Apertus.
- https://arxiv.org/abs/2406.17557 recorded 2026-07-04
FineWeb-Edu matches full FineWeb on MMLU/ARC at ~10x fewer tokens
- https://arxiv.org/abs/2502.02737 recorded 2026-07-04
SmolLM2 uses FineWeb-Edu as its English web component
Unchanged since 2026-07-04 (last edited, not re-checked)