AI Potluck
Model components / Training & synthetic datasets

FineWeb

Hugging Face

FineWeb is a large-scale English pretraining dataset of over 18.5 trillion tokens of cleaned and deduplicated text derived from 96 CommonCrawl snapshots (2013-2024), released by Hugging Face's FineWeb team. It is distributed as Parquet across ~114 date-based subsets. It was built to provide an open, high-quality web corpus that outperforms other public web datasets for LLM training, and is a widely cited reference in the open-data pretraining community.

Verified live 2026-06-22 via primary sources. Permissive ODC-BY license with ungated download and a full dataset card.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present
format
parquet

Permissive ODC-BY license with ungated download and a full dataset card.

Adoption

4 high confidence
4.0

337,208 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 high confidence
4.0

The FineWeb paper's 1.7B ablation beats C4, RefinedWeb, RedPajama, Dolma, The Pile, and OSCAR; but its FineWeb-Edu subset is the preferred quality source.

Unchanged since 2026-07-04 (last edited, not re-checked)