AI Potluck
Model components / Training & synthetic datasets

RefinedWeb

Technology Innovation Institute

English web dataset of roughly 968 million rows and about 2.8TB, built by the Technology Innovation Institute through stringent filtering and exact-plus-fuzzy deduplication of Common Crawl. The public release is an extract of the larger corpus used to train the Falcon 7B and 40B models.

Verified 2026-08-13 via the tiiuae/falcon-refinedweb dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY-1.0
gated
false
dataset_card
present
note
public_extract

Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus). The public extract is 600B tokens.

Adoption

3 high confidence
3.0

93,929 downloads in the trailing 30 days for tiiuae/falcon-refinedweb, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Filtered web alone beat The Pile and was the best legacy corpus in DCLM's tests; powered Falcon (2023-24) but is now a component inside newer pipelines rather than a standalone default.

  • https://arxiv.org/abs/2306.01116 recorded 2026-08-13

    RefinedWeb: filtered/deduplicated web alone outperforms models trained on The Pile

  • https://arxiv.org/abs/2406.11794 recorded 2026-08-13

    Abstract - DCLM-Baseline reaches 64% MMLU at 2.6T tokens, 6.6 points over the prior open-data state of the art. The legacy-corpus ranking that places RefinedWeb is in the paper body, not on this page

Verified 2026-08-13