FineWeb-2
Hugging FaceMultilingual successor to FineWeb, covering more than a thousand languages and writing systems with over a trillion tokens of filtered and deduplicated Common Crawl text. Each language runs through its own pipeline with deduplication and quality filtering, extending the FineWeb work beyond English.
Verified 2026-08-13 via the HuggingFaceFW/fineweb-2 dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- languages
- 1000+
Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-08-13
Dataset page exists on public hub, ODC-BY, 1000+ languages
Adoption
3 high confidence82,352 downloads in the trailing 30 days for HuggingFaceFW/fineweb-2, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb-2 recorded 2026-08-13
82,352 downloads in the trailing 30 days for HuggingFaceFW/fineweb-2
Capability
5 high confidencePrimary web corpus for Apertus 8B/70B (2025), 1811 languages; the current multilingual default, and its HQ variant leads multilingual data-selection work. Its card reports FineTasks wins over CulturaX, mC4, CC-100 and HPLT, and Apertus is trained on 15T tokens at 8B and 70B.
- https://arxiv.org/abs/2509.14233 recorded 2026-08-13
Abstract - Apertus 8B and 70B train on 15T tokens from over 1800 languages. The FineWeb-2 attribution is in the paper body, not on this page
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-08-13
FineWeb-2 outperforms CulturaX/mC4/CC-100/HPLT on FineTasks in controlled ablations
Verified 2026-08-13