AI Potluck
Model components / Training & synthetic datasets

Common Corpus

PleIAs

Common Corpus is a multilingual pretraining dataset of roughly 2.27 trillion tokens containing only uncopyrighted or openly licensed content across 367+ languages. It was assembled by PleIAs with partners including the AI Alliance, organized into six curated collections (public-domain books, government documents, academic content, code, semantic data, and web text) with document-level provenance and toxicity filtering. It is one of the largest openly licensed text corpora for LLM training.

Verified live 2026-06-22 via primary sources. Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.

Openness

5 high confidence
5.0
license
open(uncopyrighted/freely-licensed only)
ungated
yes
dataset_card
yes

Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.

Adoption

3 high confidence
3.0

91,700 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Trains the Pleias 1.0 SLM family and several European models (2024-25), but small demonstrator models with no controlled ablation wins; leads on ethical sourcing rather than capability.

Unchanged since 2026-07-04 (last edited, not re-checked)