AI Potluck
Model components / Training & synthetic datasets

SYNTH

PleIAs

SYNTH is a large-scale synthetic dataset from PleIAs (released with the AI Alliance) designed for training small reasoning models end-to-end. It contains tens of millions of synthetic samples amplifying content from tens of thousands of Wikipedia articles, Wikibooks, and internal documents, spanning exercise types like arithmetic, creative writing, RAG, and information extraction. Samples include reasoning traces and seed attribution, and roughly 20% covers non-English European languages.

Verified live 2026-06-22 via primary sources. Permissively licensed, ungated dataset with a documented card.

Openness

5 high confidence
5.0
license
CC-BY-4.0(permissive)

Permissively licensed, ungated dataset with a documented card.

Adoption

2 high confidence
2.0

6,105 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 medium confidence
4.0

A novel fully-open synthetic reasoning corpus with real first-party end-to-end results (Baguettotron best-in-class sub-350M), but single-lab evidence.

Unchanged since 2026-07-04 (last edited, not re-checked)