AI Potluck
Model components / Training & synthetic datasets

OSCAR-2301

OSCAR Project

OSCAR-2301 is a multilingual web corpus extracted and filtered from the late-2022 Common Crawl snapshot, covering 151 languages and totaling ~3.57TB of raw unannotated text for LLM pretraining. It is published by the OSCAR project (oscar-corpus) and includes KenLM-based content quality metrics and deduplication hashes. The corpus content is from Common Crawl with no copyright held by the authors; metadata is CC0-1.0.

Verified live 2026-06-22 via primary sources. Manually gated and the card states access is temporarily suspended pending clarification.

Openness

3 high confidence
3.0
license
cc0-1.0(metadata)
gated
manual(login + terms acceptance)

Manually gated and the card states access is temporarily suspended pending clarification.

Adoption

2 high confidence
2.0

1,793 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

A dated 2022-snapshot web corpus with access suspended; the derived Colossal OSCAR trains ALIA/Salamandra (2024) but the 2301 version itself is superseded by FineWeb-2/CulturaX.

Unchanged since 2026-07-04 (last edited, not re-checked)