RefinedWeb
Technology Innovation InstituteEnglish web dataset of roughly 968 million rows and about 2.8TB, built by the Technology Innovation Institute through stringent filtering and exact-plus-fuzzy deduplication of Common Crawl. The public release is an extract of the larger corpus used to train the Falcon 7B and 40B models.
Verified 2026-08-13 via the tiiuae/falcon-refinedweb dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY-1.0
- gated
- false
- dataset_card
- present
- note
- public_extract
Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus). The public extract is 600B tokens.
- https://huggingface.co/datasets/tiiuae/falcon-refinedweb recorded 2026-08-13
Dataset card, ODC-By 1.0 license, ungated public download
Adoption
3 high confidence93,929 downloads in the trailing 30 days for tiiuae/falcon-refinedweb, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/tiiuae/falcon-refinedweb recorded 2026-08-13
93,929 downloads in the trailing 30 days for tiiuae/falcon-refinedweb
Capability
3 high confidenceFiltered web alone beat The Pile and was the best legacy corpus in DCLM's tests; powered Falcon (2023-24) but is now a component inside newer pipelines rather than a standalone default.
- https://arxiv.org/abs/2306.01116 recorded 2026-08-13
RefinedWeb: filtered/deduplicated web alone outperforms models trained on The Pile
- https://arxiv.org/abs/2406.11794 recorded 2026-08-13
Abstract - DCLM-Baseline reaches 64% MMLU at 2.6T tokens, 6.6 points over the prior open-data state of the art. The legacy-corpus ranking that places RefinedWeb is in the paper body, not on this page
Verified 2026-08-13