RefinedWeb
Technology Innovation InstituteFalcon RefinedWeb is a large English web dataset (~968M rows, ~2.8TB) built by the Technology Innovation Institute (TII) via stringent filtering and exact-plus-fuzzy deduplication of CommonCrawl. The public release is an extract of the larger corpus used to train the Falcon-7B/40B models. It demonstrated that rigorously filtered web data alone can match or beat curated corpora for LLM pretraining.
Verified live 2026-06-22 via primary sources. Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus).
Openness
5 high confidence- license
- ODC-BY-1.0
- gated
- false
- dataset_card
- present
- note
- public_extract
Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus).
- https://huggingface.co/datasets/tiiuae/falcon-refinedweb recorded 2026-06-22
Dataset card, ODC-By 1.0 license, ungated public download
Adoption
3 high confidence12,934 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/tiiuae/falcon-refinedweb recorded 2026-07-04
downloads field = 12,934 (30-day)
Capability
3 high confidenceFiltered web alone beat The Pile and was the best legacy corpus in DCLM's tests; powered Falcon (2023-24) but is now a component inside newer pipelines rather than a standalone default.
- https://arxiv.org/abs/2306.01116 recorded 2026-07-04
RefinedWeb: filtered/deduplicated web alone outperforms models trained on The Pile
- https://arxiv.org/abs/2406.11794 recorded 2026-07-04
In DCLM's controlled ablations RefinedWeb is the best legacy corpus but DCLM-Baseline surpasses it
Unchanged since 2026-07-04 (last edited, not re-checked)