AI Potluck
Model components / Training & synthetic datasets

CulturaX

University of Oregon NLP (UONLP)

Multilingual web corpus for pretraining containing 6.3 trillion tokens across 167 languages, combined and cleaned from mC4 and OSCAR by the University of Oregon NLP group. Its pipeline applies language identification, URL filtering, metric-based cleaning, document refinement and MinHash fuzzy deduplication, and it is distributed as roughly 16TB of parquet.

Verified 2026-08-13 via the uonlp/CulturaX dataset card on Hugging Face.

Openness

3 high confidence
3.0
license
follows mC4 + OSCAR-2301 terms
gated
auto(contact-info agreement)
card
present

Dataset card is rich but downloads require agreeing to share contact information (auto gate).

Adoption

3 high confidence
3.0

13,746 downloads in the trailing 30 days for uonlp/CulturaX, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

3 high confidence
3.0

A strong 2023-24 default multilingual corpus, but explicitly beaten by FineWeb-2 in controlled per-language ablations. CulturaX is a 6.3T-token, 167-language corpus built from mC4 and OSCAR.

Verified 2026-08-13