Common Corpus
PleIAsMultilingual pretraining dataset of roughly 2.27 trillion tokens across more than 367 languages, assembled by PleIAs with partners including the AI Alliance from uncopyrighted or openly licensed content only. It is organized into six curated collections - public-domain books, government documents, academic content, code, semantic data and web text - with document-level provenance and toxicity filtering.
Verified 2026-08-13 via the PleIAs/common_corpus dataset card on Hugging Face.
Openness
5 high confidence- license
- open(uncopyrighted/freely-licensed only)
- ungated
- yes
- dataset_card
- yes
Ungated with a full dataset card and built exclusively from uncopyrighted or openly licensed sources.
- https://huggingface.co/datasets/PleIAs/common_corpus recorded 2026-08-13
open-license-only content, ungated access, dataset card, 2.27T tokens
Adoption
3 high confidence61,684 downloads in the trailing 30 days for PleIAs/common_corpus, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/PleIAs/common_corpus recorded 2026-08-13
61,684 downloads in the trailing 30 days for PleIAs/common_corpus
Capability
3 high confidenceTrains the Pleias 1.0 SLM family and several European models (2024-25), but small demonstrator models with no controlled ablation wins; leads on ethical sourcing rather than capability. The corpus is 2.27 trillion tokens.
- https://arxiv.org/html/2506.01732v1 recorded 2026-08-13
Pleias 1.0 family and ~7 European models trained on Common Corpus
- https://huggingface.co/blog/Pclanglais/common-corpus recorded 2026-08-13
Common Corpus presented as the largest open licensed pretraining corpus; the 2.27T-token figure is on the dataset card
Verified 2026-08-13