AI Potluck
Model components / Training & synthetic datasets

RefinedWeb

Technology Innovation Institute

Falcon RefinedWeb is a large English web dataset (~968M rows, ~2.8TB) built by the Technology Innovation Institute (TII) via stringent filtering and exact-plus-fuzzy deduplication of CommonCrawl. The public release is an extract of the larger corpus used to train the Falcon-7B/40B models. It demonstrated that rigorously filtered web data alone can match or beat curated corpora for LLM pretraining.

Verified live 2026-06-22 via primary sources. Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus).

Openness

5 high confidence
5.0
license
ODC-BY-1.0
gated
false
dataset_card
present
note
public_extract

Permissive ODC-BY 1.0 license, ungated download, with a dataset card (a public extract of the full corpus).

Adoption

3 high confidence
3.0

12,934 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Filtered web alone beat The Pile and was the best legacy corpus in DCLM's tests; powered Falcon (2023-24) but is now a component inside newer pipelines rather than a standalone default.

Unchanged since 2026-07-04 (last edited, not re-checked)