SYNTH
PleIAsSYNTH is a large-scale synthetic dataset from PleIAs (released with the AI Alliance) designed for training small reasoning models end-to-end. It contains tens of millions of synthetic samples amplifying content from tens of thousands of Wikipedia articles, Wikibooks, and internal documents, spanning exercise types like arithmetic, creative writing, RAG, and information extraction. Samples include reasoning traces and seed attribution, and roughly 20% covers non-English European languages.
Verified live 2026-06-22 via primary sources. Permissively licensed, ungated dataset with a documented card.
Openness
5 high confidence- license
- CC-BY-4.0(permissive)
Permissively licensed, ungated dataset with a documented card.
- https://huggingface.co/datasets/PleIAs/SYNTH recorded 2026-06-22
License CC-BY-4.0, not gated, dataset card present
Adoption
2 high confidence6,105 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/PleIAs/SYNTH recorded 2026-07-04
downloads field = 6,105 (30-day)
Capability
4 medium confidenceA novel fully-open synthetic reasoning corpus with real first-party end-to-end results (Baguettotron best-in-class sub-350M), but single-lab evidence.
- https://huggingface.co/blog/Pclanglais/synth-data-frontier recorded 2026-07-04
PleIAs trained Baguettotron (321M) and Monad (56M) end-to-end on SYNTH, best-in-class for size
- https://huggingface.co/datasets/PleIAs/SYNTH recorded 2026-07-04
SYNTH: ~79.6M samples, CC-BY-4.0, released Nov 2025
Unchanged since 2026-07-04 (last edited, not re-checked)