FineWeb
Hugging FaceFineWeb is a large-scale English pretraining dataset of over 18.5 trillion tokens of cleaned and deduplicated text derived from 96 CommonCrawl snapshots (2013-2024), released by Hugging Face's FineWeb team. It is distributed as Parquet across ~114 date-based subsets. It was built to provide an open, high-quality web corpus that outperforms other public web datasets for LLM training, and is a widely cited reference in the open-data pretraining community.
Verified live 2026-06-22 via primary sources. Permissive ODC-BY license with ungated download and a full dataset card.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- format
- parquet
Permissive ODC-BY license with ungated download and a full dataset card.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb recorded 2026-06-22
Dataset card present, ODC-BY license, public ungated download
Adoption
4 high confidence337,208 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb recorded 2026-07-04
downloads field = 337,208 (30-day)
Capability
4 high confidenceThe FineWeb paper's 1.7B ablation beats C4, RefinedWeb, RedPajama, Dolma, The Pile, and OSCAR; but its FineWeb-Edu subset is the preferred quality source.
- https://arxiv.org/abs/2406.17557 recorded 2026-07-04
FineWeb 1.7B ablation outperforms C4/RefinedWeb/RedPajama/Dolma/Pile/OSCAR on the aggregate suite
Unchanged since 2026-07-04 (last edited, not re-checked)