AI Potluck
Model components / Training & synthetic datasets

Cosmopedia

Hugging Face

Synthetic pretraining dataset of more than 31 million textbooks, blog posts, stories and educational articles, about 25 billion tokens, generated with Mixtral-8x7B-Instruct from seed samples drawn from web data, Stanford courses, OpenStax, Khan Academy and WikiHow. It was built by Hugging Face's Smol Models Research team as a reference corpus for synthetic-data pretraining recipes.

Cosmopedia v2 is the direct successor. Verified 2026-08-13 via the HuggingFaceTB/cosmopedia dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
apache-2.0
access
ungated
dataset_card
present

Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.

Adoption

3 high confidence
3.0

22,057 downloads in the trailing 30 days for HuggingFaceTB/cosmopedia, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Cosmo-1B beats TinyLlama on ARC/MMLU but trails Phi-1.5; explicitly superseded by Cosmopedia v2.

Verified 2026-08-13