AI Potluck
Model components / Training & synthetic datasets

FineWeb-2

Hugging Face

Multilingual successor to FineWeb, covering more than a thousand languages and writing systems with over a trillion tokens of filtered and deduplicated Common Crawl text. Each language runs through its own pipeline with deduplication and quality filtering, extending the FineWeb work beyond English.

Verified 2026-08-13 via the HuggingFaceFW/fineweb-2 dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present
languages
1000+

Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.

Adoption

3 high confidence
3.0

82,352 downloads in the trailing 30 days for HuggingFaceFW/fineweb-2, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

5 high confidence
5.0

Primary web corpus for Apertus 8B/70B (2025), 1811 languages; the current multilingual default, and its HQ variant leads multilingual data-selection work. Its card reports FineTasks wins over CulturaX, mC4, CC-100 and HPLT, and Apertus is trained on 15T tokens at 8B and 70B.

Verified 2026-08-13