Common Corpus
PleIAsCommon Corpus is a multilingual pretraining dataset of roughly 2.27 trillion tokens containing only uncopyrighted or openly licensed content across 367+ languages. It was assembled by PleIAs with partners including the AI Alliance, organized into six curated collections (public-domain books, government documents, academic content, code, semantic data, and web text) with document-level provenance and toxicity filtering. It is one of the largest openly licensed text corpora for LLM training.
Verified live 2026-06-22 via primary sources. Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.
Openness
5 high confidence- license
- open(uncopyrighted/freely-licensed only)
- ungated
- yes
- dataset_card
- yes
Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.
- https://huggingface.co/datasets/PleIAs/common_corpus recorded 2026-06-22
open-license-only content, ungated access, dataset card, 2.27T tokens
Adoption
3 high confidence91,700 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/PleIAs/common_corpus recorded 2026-07-04
downloads field = 91,700 (30-day)
Capability
3 high confidenceTrains the Pleias 1.0 SLM family and several European models (2024-25), but small demonstrator models with no controlled ablation wins; leads on ethical sourcing rather than capability.
- https://arxiv.org/html/2506.01732v1 recorded 2026-07-04
Pleias 1.0 family and ~7 European models trained on Common Corpus
- https://huggingface.co/blog/Pclanglais/common-corpus recorded 2026-07-04
Common Corpus is the largest fully-open licensed pretraining corpus (2.27T tokens)
Unchanged since 2026-07-04 (last edited, not re-checked)