Cosmopedia
Hugging FaceSynthetic pretraining dataset of more than 31 million textbooks, blog posts, stories and educational articles, about 25 billion tokens, generated with Mixtral-8x7B-Instruct from seed samples drawn from web data, Stanford courses, OpenStax, Khan Academy and WikiHow. It was built by Hugging Face's Smol Models Research team as a reference corpus for synthetic-data pretraining recipes.
Cosmopedia v2 is the direct successor. Verified 2026-08-13 via the HuggingFaceTB/cosmopedia dataset card on Hugging Face.
Openness
5 high confidence- license
- apache-2.0
- access
- ungated
- dataset_card
- present
Permissive Apache-2.0 license, ungated, full dataset card and downloadable parquet files.
- https://huggingface.co/datasets/HuggingFaceTB/cosmopedia recorded 2026-08-13
Apache-2.0 license, 31M rows, full dataset card
Adoption
3 high confidence22,057 downloads in the trailing 30 days for HuggingFaceTB/cosmopedia, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/HuggingFaceTB/cosmopedia recorded 2026-08-13
22,057 downloads in the trailing 30 days for HuggingFaceTB/cosmopedia
Capability
3 high confidenceCosmo-1B beats TinyLlama on ARC/MMLU but trails Phi-1.5; explicitly superseded by Cosmopedia v2.
- https://huggingface.co/blog/cosmopedia recorded 2026-08-13
Cosmo-1B performs better than TinyLlama 1.1B on ARC-easy, ARC-challenge, OpenBookQA and MMLU, with gaps remaining against Phi-1.5
- https://huggingface.co/datasets/HuggingFaceTB/cosmopedia-v2 recorded 2026-08-13
Cosmopedia v2 is the direct successor
Verified 2026-08-13