AI Potluck
Model components / Training & synthetic datasets

Cosmopedia

Hugging Face

Cosmopedia is a large-scale synthetic dataset of over 31 million textbooks, blog posts, stories, and educational articles (~25 billion tokens) generated with Mixtral-8x7B-Instruct using seed samples from web data, Stanford courses, OpenStax, Khan Academy, and WikiHow. It was created by Hugging Face's Smol Models Research team and described as the largest open synthetic dataset of its kind at release. It is a reference corpus for synthetic-data pretraining recipes.

Verified live 2026-06-22 via primary sources. Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.

Openness

5 high confidence
5.0
license
apache-2.0

Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.

Adoption

3 high confidence
3.0

16,870 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Cosmo-1B beats TinyLlama on ARC/MMLU but trails Phi-1.5; explicitly superseded by Cosmopedia v2.

Unchanged since 2026-07-04 (last edited, not re-checked)