AI Potluck
Model components / Training & synthetic datasets

FineWeb-Edu

Hugging Face

Educational subset of FineWeb holding about 1.3 trillion tokens, produced by filtering FineWeb with an educational-quality classifier trained on Llama-3-70B annotations and keeping the text scored as highly educational. It targets knowledge- and reasoning-heavy evaluation relative to unfiltered web data.

Verified 2026-08-13 via the HuggingFaceFW/fineweb-edu dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present
format
parquet

Permissive ODC-BY license with ungated download and full dataset card.

Adoption

4 high confidence
4.0

395,837 downloads in the trailing 30 days for HuggingFaceFW/fineweb-edu, which bands at 100K-1M, level 4 on the dataset adoption scale.

Capability

5 high confidence
5.0

Dominates FineWeb and all peers on MMLU/ARC, matching full FineWeb with ~10x fewer tokens; current default web corpus for SmolLM2/3 and Apertus.

  • https://arxiv.org/abs/2406.17557 recorded 2026-08-13

    Abstract - LLMs pretrained on FineWeb-Edu, a 1.3T-token educational subset, do dramatically better on MMLU and ARC. The ~10x token-efficiency figure is in the paper body, not on this page

  • https://arxiv.org/abs/2502.02737 recorded 2026-08-13

    Abstract - SmolLM2 is overtrained on ~11T tokens mixing web text with specialised math, code and instruction data. The named FineWeb-Edu component is in the paper body, not on this page

Verified 2026-08-13