AI Potluck
Model components / Training & synthetic datasets

FineWeb-2

Hugging Face

FineWeb-2 is the multilingual successor to FineWeb, covering over 1,000 languages and writing systems with more than 1 trillion tokens of filtered and deduplicated CommonCrawl text, released by Hugging Face's FineWeb team. Each language is processed through a per-language pipeline with deduplication and quality filtering. It extends the open-data pretraining work beyond English to support multilingual LLM training.

Verified live 2026-06-22 via primary sources. Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present
languages
1000+

Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.

Adoption

3 high confidence
3.0

93,906 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

5 high confidence
5.0

Primary web corpus for Apertus 8B/70B (2025), 1811 languages; the current multilingual default, and its HQ variant leads multilingual data-selection work.

Unchanged since 2026-07-04 (last edited, not re-checked)