Cosmopedia
Hugging FaceCosmopedia is a large-scale synthetic dataset of over 31 million textbooks, blog posts, stories, and educational articles (~25 billion tokens) generated with Mixtral-8x7B-Instruct using seed samples from web data, Stanford courses, OpenStax, Khan Academy, and WikiHow. It was created by Hugging Face's Smol Models Research team and described as the largest open synthetic dataset of its kind at release. It is a reference corpus for synthetic-data pretraining recipes.
Verified live 2026-06-22 via primary sources. Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.
Openness
5 high confidence- license
- apache-2.0
Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.
- https://huggingface.co/datasets/HuggingFaceTB/cosmopedia recorded 2026-06-22
Apache-2.0 license, 31M rows, full dataset card
Adoption
3 high confidence16,870 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/HuggingFaceTB/cosmopedia recorded 2026-07-04
downloads field = 16,870 (30-day)
Capability
3 high confidenceCosmo-1B beats TinyLlama on ARC/MMLU but trails Phi-1.5; explicitly superseded by Cosmopedia v2.
- https://huggingface.co/blog/cosmopedia recorded 2026-07-04
Cosmo-1B beats TinyLlama on ARC/OpenBookQA/MMLU, still trails Phi-1.5
- https://huggingface.co/datasets/HuggingFaceTB/cosmopedia-v2 recorded 2026-07-04
Cosmopedia v2 is the direct successor
Unchanged since 2026-07-04 (last edited, not re-checked)