AI Potluck
Model components / Training & synthetic datasets

FineWeb

Hugging Face

English pretraining dataset of more than 18.5 trillion tokens of cleaned and deduplicated text derived from 96 Common Crawl snapshots between 2013 and 2024, released by Hugging Face's FineWeb team. It is distributed as Parquet across roughly 114 date-based subsets.

Verified 2026-08-13 via the HuggingFaceFW/fineweb dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present
format
parquet

Permissive ODC-BY license with ungated download and a full dataset card.

Adoption

4 high confidence
4.0

418,578 downloads in the trailing 30 days for HuggingFaceFW/fineweb, which bands at 100K-1M, level 4 on the dataset adoption scale.

Capability

4 high confidence
4.0

The FineWeb paper's 1.7B ablation beats C4, RefinedWeb, RedPajama, Dolma, The Pile, and OSCAR; but its FineWeb-Edu subset is the preferred quality source.

  • https://arxiv.org/abs/2406.17557 recorded 2026-08-13

    Abstract - FineWeb, 15T tokens from 96 Common Crawl snapshots, produces better-performing LLMs than other open pretraining datasets. The per-corpus ablation table is in the paper body, not on this page

Verified 2026-08-13