AI Potluck
Model components / Training & synthetic datasets

SYNTH

PleIAs

Large synthetic dataset from PleIAs, released with the AI Alliance, designed for training small reasoning models end to end. It holds tens of millions of samples amplifying content from tens of thousands of Wikipedia articles, Wikibooks and internal documents, across exercise types including arithmetic, creative writing, RAG and information extraction. Samples carry reasoning traces and seed attribution, and roughly a fifth covers non-English European languages.

PleIAs trained Baguettotron and Monad end to end on it. Verified 2026-08-13 via the PleIAs/SYNTH dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
CC-BY-4.0(permissive)
access
ungated
dataset_card
present

Permissively licensed, ungated dataset with a documented card.

Adoption

3 medium confidence
3.0

The Hugging Face datasets API reports 10,111 downloads in the trailing 30 days for PleIAs/SYNTH, ungated and with 273 likes, which bands at 10K-100K, level 3 on the dataset adoption scale. This is a near-boundary reading and should be read as one: 10,111 clears the 10,000 floor of the band by 1.1%, so a month of ordinary variation could put it back at level 2. Confidence is medium for that reason rather than because of anything wrong with the measurement.

Capability

4 medium confidence
4.0

A novel fully-open synthetic reasoning corpus with real first-party end-to-end results (Baguettotron best-in-class sub-350M), but single-lab evidence. The PleIAs post reports Baguettotron at 321M and Monad at 56M, both trained end to end on SYNTH.

Verified 2026-08-13