FineWeb
Hugging FaceEnglish pretraining dataset of more than 18.5 trillion tokens of cleaned and deduplicated text derived from 96 Common Crawl snapshots between 2013 and 2024, released by Hugging Face's FineWeb team. It is distributed as Parquet across roughly 114 date-based subsets.
Verified 2026-08-13 via the HuggingFaceFW/fineweb dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- format
- parquet
Permissive ODC-BY license with ungated download and a full dataset card.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb recorded 2026-08-13
Dataset card present, ODC-BY license, public ungated download
Adoption
4 high confidence418,578 downloads in the trailing 30 days for HuggingFaceFW/fineweb, which bands at 100K-1M, level 4 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb recorded 2026-08-13
418,578 downloads in the trailing 30 days for HuggingFaceFW/fineweb
Capability
4 high confidenceThe FineWeb paper's 1.7B ablation beats C4, RefinedWeb, RedPajama, Dolma, The Pile, and OSCAR; but its FineWeb-Edu subset is the preferred quality source.
- https://arxiv.org/abs/2406.17557 recorded 2026-08-13
Abstract - FineWeb, 15T tokens from 96 Common Crawl snapshots, produces better-performing LLMs than other open pretraining datasets. The per-corpus ablation table is in the paper body, not on this page
Verified 2026-08-13