SYNTH
PleIAsLarge synthetic dataset from PleIAs, released with the AI Alliance, designed for training small reasoning models end to end. It holds tens of millions of samples amplifying content from tens of thousands of Wikipedia articles, Wikibooks and internal documents, across exercise types including arithmetic, creative writing, RAG and information extraction. Samples carry reasoning traces and seed attribution, and roughly a fifth covers non-English European languages.
PleIAs trained Baguettotron and Monad end to end on it. Verified 2026-08-13 via the PleIAs/SYNTH dataset card on Hugging Face.
Openness
5 high confidence- license
- CC-BY-4.0(permissive)
- access
- ungated
- dataset_card
- present
Permissively licensed, ungated dataset with a documented card.
- https://huggingface.co/datasets/PleIAs/SYNTH recorded 2026-08-13
License CC-BY-4.0, not gated, dataset card present
Adoption
3 medium confidenceThe Hugging Face datasets API reports 10,111 downloads in the trailing 30 days for PleIAs/SYNTH, ungated and with 273 likes, which bands at 10K-100K, level 3 on the dataset adoption scale. This is a near-boundary reading and should be read as one: 10,111 clears the 10,000 floor of the band by 1.1%, so a month of ordinary variation could put it back at level 2. Confidence is medium for that reason rather than because of anything wrong with the measurement.
- https://huggingface.co/api/datasets/PleIAs/SYNTH recorded 2026-08-14
downloads field = 10,111 in the trailing 30 days; 273 likes; gated false
Capability
4 medium confidenceA novel fully-open synthetic reasoning corpus with real first-party end-to-end results (Baguettotron best-in-class sub-350M), but single-lab evidence. The PleIAs post reports Baguettotron at 321M and Monad at 56M, both trained end to end on SYNTH.
- https://huggingface.co/blog/Pclanglais/synth-data-frontier recorded 2026-08-13
PleIAs trained Baguettotron (321M) and Monad (56M) end-to-end on SYNTH, best-in-class for size
- https://huggingface.co/datasets/PleIAs/SYNTH recorded 2026-08-13
SYNTH: ~79.6M samples, CC-BY-4.0, released Nov 2025
Verified 2026-08-13