AI Potluck
Model components / Training & synthetic datasets

Common Corpus

PleIAs

Multilingual pretraining dataset of roughly 2.27 trillion tokens across more than 367 languages, assembled by PleIAs with partners including the AI Alliance from uncopyrighted or openly licensed content only. It is organized into six curated collections - public-domain books, government documents, academic content, code, semantic data and web text - with document-level provenance and toxicity filtering.

Verified 2026-08-13 via the PleIAs/common_corpus dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
open(uncopyrighted/freely-licensed only)
ungated
yes
dataset_card
yes

Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.

Adoption

3 high confidence
3.0

61,684 downloads in the trailing 30 days for PleIAs/common_corpus, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

Trains the Pleias 1.0 SLM family and several European models (2024-25), but small demonstrator models with no controlled ablation wins; leads on ethical sourcing rather than capability. The corpus is 2.27 trillion tokens.

Verified 2026-08-13