FineWeb-Edu
Hugging FaceEducational subset of FineWeb holding about 1.3 trillion tokens, produced by filtering FineWeb with an educational-quality classifier trained on Llama-3-70B annotations and keeping the text scored as highly educational. It targets knowledge- and reasoning-heavy evaluation relative to unfiltered web data.
Verified 2026-08-13 via the HuggingFaceFW/fineweb-edu dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- format
- parquet
Permissive ODC-BY license with ungated download and full dataset card.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu recorded 2026-08-13
Dataset card present, ODC-BY license, public ungated download
Adoption
4 high confidence395,837 downloads in the trailing 30 days for HuggingFaceFW/fineweb-edu, which bands at 100K-1M, level 4 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb-edu recorded 2026-08-13
395,837 downloads in the trailing 30 days for HuggingFaceFW/fineweb-edu
Capability
5 high confidenceDominates FineWeb and all peers on MMLU/ARC, matching full FineWeb with ~10x fewer tokens; current default web corpus for SmolLM2/3 and Apertus.
- https://arxiv.org/abs/2406.17557 recorded 2026-08-13
Abstract - LLMs pretrained on FineWeb-Edu, a 1.3T-token educational subset, do dramatically better on MMLU and ARC. The ~10x token-efficiency figure is in the paper body, not on this page
- https://arxiv.org/abs/2502.02737 recorded 2026-08-13
Abstract - SmolLM2 is overtrained on ~11T tokens mixing web text with specialised math, code and instruction data. The named FineWeb-Edu component is in the paper body, not on this page
Verified 2026-08-13