AI Potluck
Model components / Training & synthetic datasets

C4

Allen Institute for AI

Cleaned version of Common Crawl created for the T5 paper and hosted by AI2. It ships several variants, including a cleaned English split of about 305GB and the multilingual mC4 covering 108 languages at roughly 9.7TB.

Verified 2026-08-13 via the allenai/c4 dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
ungated
yes
dataset_card
yes

Permissively licensed (ODC-BY) and ungated with a full dataset card.

Adoption

5 high confidence
5.0

1,414,303 downloads in the trailing 30 days for allenai/c4, which bands at >1M, level 5 on the dataset adoption scale.

Capability

3 high confidence
3.0

Powered T5 (2019) and UL2 (2022); now a baseline that FineWeb and DCLM beat, superseded by modern filtered web corpora.

  • https://arxiv.org/abs/2406.17557 recorded 2026-08-13

    Abstract - FineWeb produces better-performing LLMs than other open pretraining datasets. The C4 baseline row is in the paper body, not on this page

Verified 2026-08-13