CulturaX
University of Oregon NLP (UONLP)CulturaX is a multilingual web corpus for LLM pretraining containing 6.3 trillion tokens across 167 languages, combined and cleaned from mC4 and OSCAR. It was created by the University of Oregon NLP group (UONLP) and applies a multi-stage pipeline of language identification, URL filtering, metric-based cleaning, document refinement, and MinHash fuzzy deduplication. Distributed as ~16TB of parquet.
Verified live 2026-06-22 via primary sources. Dataset card is rich but downloads require agreeing to share contact information (auto gate).
Openness
3 high confidence- license
- follows mC4 + OSCAR-2301 terms
- gated
- auto(contact-info agreement)
- card
- present
Dataset card is rich but downloads require agreeing to share contact information (auto gate).
- https://huggingface.co/datasets/uonlp/CulturaX recorded 2026-06-22
Card with 6.3T tokens/167 languages, mC4+OSCAR license inheritance, gated access requiring contact-info agreement
Adoption
3 high confidence18,361 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/uonlp/CulturaX recorded 2026-07-04
downloads field = 18,361 (30-day)
Capability
3 high confidenceA strong 2023-24 default multilingual corpus, but explicitly beaten by FineWeb-2 in controlled per-language ablations.
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-07-04
FineWeb-2 outperforms CulturaX across diverse languages on FineTasks
- https://huggingface.co/datasets/uonlp/CulturaX recorded 2026-07-04
CulturaX is a 6.3T-token, 167-language corpus built from mC4+OSCAR
Unchanged since 2026-07-04 (last edited, not re-checked)