FineWeb-2
Hugging FaceFineWeb-2 is the multilingual successor to FineWeb, covering over 1,000 languages and writing systems with more than 1 trillion tokens of filtered and deduplicated CommonCrawl text, released by Hugging Face's FineWeb team. Each language is processed through a per-language pipeline with deduplication and quality filtering. It extends the open-data pretraining work beyond English to support multilingual LLM training.
Verified live 2026-06-22 via primary sources. Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
- languages
- 1000+
Permissive ODC-BY license, ungated, with a dataset card spanning 1000+ languages.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-06-22
Dataset page exists on public hub, ODC-BY, 1000+ languages
Adoption
3 high confidence93,906 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/HuggingFaceFW/fineweb-2 recorded 2026-07-04
downloads field = 93,906 (30-day)
Capability
5 high confidencePrimary web corpus for Apertus 8B/70B (2025), 1811 languages; the current multilingual default, and its HQ variant leads multilingual data-selection work.
- https://arxiv.org/abs/2509.14233 recorded 2026-07-04
Apertus 8B/70B train on 15T tokens from 1811 languages via FineWeb-2/FineWeb-2-HQ
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-07-04
FineWeb-2 outperforms CulturaX/mC4/CC-100/HPLT on FineTasks in controlled ablations
Unchanged since 2026-07-04 (last edited, not re-checked)