AI Potluck
Model components / Training & synthetic datasets

CulturaX

University of Oregon NLP (UONLP)

CulturaX is a multilingual web corpus for LLM pretraining containing 6.3 trillion tokens across 167 languages, combined and cleaned from mC4 and OSCAR. It was created by the University of Oregon NLP group (UONLP) and applies a multi-stage pipeline of language identification, URL filtering, metric-based cleaning, document refinement, and MinHash fuzzy deduplication. Distributed as ~16TB of parquet.

Verified live 2026-06-22 via primary sources. Dataset card is rich but downloads require agreeing to share contact information (auto gate).

Openness

3 high confidence
3.0
license
follows mC4 + OSCAR-2301 terms
gated
auto(contact-info agreement)
card
present

Dataset card is rich but downloads require agreeing to share contact information (auto gate).

Adoption

3 high confidence
3.0

18,361 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

A strong 2023-24 default multilingual corpus, but explicitly beaten by FineWeb-2 in controlled per-language ablations.

Unchanged since 2026-07-04 (last edited, not re-checked)