C4
Allen Institute for AICleaned version of Common Crawl created for the T5 paper and hosted by AI2. It ships several variants, including a cleaned English split of about 305GB and the multilingual mC4 covering 108 languages at roughly 9.7TB.
Verified 2026-08-13 via the allenai/c4 dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- ungated
- yes
- dataset_card
- yes
Permissively licensed (ODC-BY) and ungated with a full dataset card.
- https://huggingface.co/datasets/allenai/c4 recorded 2026-08-13
ODC-BY license, ungated access, dataset card, size variants
Adoption
5 high confidence1,414,303 downloads in the trailing 30 days for allenai/c4, which bands at >1M, level 5 on the dataset adoption scale.
- https://huggingface.co/api/datasets/allenai/c4 recorded 2026-08-13
1,414,303 downloads in the trailing 30 days for allenai/c4
Capability
3 high confidencePowered T5 (2019) and UL2 (2022); now a baseline that FineWeb and DCLM beat, superseded by modern filtered web corpora.
- https://arxiv.org/abs/2406.17557 recorded 2026-08-13
Abstract - FineWeb produces better-performing LLMs than other open pretraining datasets. The C4 baseline row is in the paper body, not on this page
Verified 2026-08-13