AI Potluck
Model components / Training & synthetic datasets

DCLM-baseline

DataComp-LM (ML Foundations)

DCLM-baseline-1.0 is a curated CommonCrawl-derived pretraining dataset of about 4 trillion tokens across 3 billion documents (~7.2TB), produced by the DataComp-LM (ML Foundations) project. It was built using heuristic cleaning, Bloom-filter deduplication, and fastText model-based quality filtering, and is the reference dataset for the DCLM benchmark. It is positioned as a research dataset that outperforms prior open web datasets when training 7B models.

Verified live 2026-06-22 via primary sources. Permissive CC-BY-4.0 license, ungated download, with open code repo and dataset card.

Openness

5 high confidence
5.0
license
CC-BY-4.0
gated
false
dataset_card
present
github
mlfoundations/dclm

Permissive CC-BY-4.0 license, ungated download, with open code repo and dataset card.

Adoption

4 high confidence
4.0

440,189 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

5 high confidence
5.0

Leads the DataComp-LM leaderboard (beats C4/RefinedWeb/RedPajama/Dolma); current 2024-26 default of OLMoE, Marin, SmolLM3, and RWKV.

Unchanged since 2026-07-04 (last edited, not re-checked)