AI Potluck
Model components / Training & synthetic datasets

C4

Allen Institute for AI

C4 (Colossal Clean Crawled Corpus) is a cleaned version of Common Crawl created for the T5 paper, hosted on Hugging Face by AI2. It offers multiple variants, including the cleaned English split (~305GB) and the multilingual mC4 covering 108 languages (~9.7TB). It is one of the most widely used English pretraining corpora.

Verified live 2026-06-22 via primary sources. Permissively licensed (ODC-BY) and ungated with a full dataset card.

Openness

5 high confidence
5.0
license
ODC-BY
ungated
yes
dataset_card
yes

Permissively licensed (ODC-BY) and ungated with a full dataset card.

Adoption

5 high confidence
5.0

1,261,777 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Powered T5 (2019) and UL2 (2022); now a baseline that FineWeb and DCLM beat, superseded by modern filtered web corpora.

Unchanged since 2026-07-04 (last edited, not re-checked)