DCLM-baseline
DataComp-LM (ML Foundations)DCLM-baseline-1.0 is a curated CommonCrawl-derived pretraining dataset of about 4 trillion tokens across 3 billion documents (~7.2TB), produced by the DataComp-LM (ML Foundations) project. It was built using heuristic cleaning, Bloom-filter deduplication, and fastText model-based quality filtering, and is the reference dataset for the DCLM benchmark. It is positioned as a research dataset that outperforms prior open web datasets when training 7B models.
Verified live 2026-06-22 via primary sources. Permissive CC-BY-4.0 license, ungated download, with open code repo and dataset card.
Openness
5 high confidence- license
- CC-BY-4.0
- gated
- false
- dataset_card
- present
- github
- mlfoundations/dclm
Permissive CC-BY-4.0 license, ungated download, with open code repo and dataset card.
- https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0 recorded 2026-06-22
Dataset card, CC-BY-4.0 license, ungated public download, GitHub link
Adoption
4 high confidence440,189 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/mlfoundations/dclm-baseline-1.0 recorded 2026-07-04
downloads field = 440,189 (30-day)
Capability
5 high confidenceLeads the DataComp-LM leaderboard (beats C4/RefinedWeb/RedPajama/Dolma); current 2024-26 default of OLMoE, Marin, SmolLM3, and RWKV.
- https://arxiv.org/abs/2406.11794 recorded 2026-07-04
DCLM-Baseline 7B reaches 64% MMLU at 2.6T tokens; model-based filtering adds >6 MMLU over the best prior open corpus
- https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0 recorded 2026-07-04
OLMoE, SmolLM3, Marin, and RWKV built on DCLM
Unchanged since 2026-07-04 (last edited, not re-checked)