OSCAR-2301
OSCAR ProjectOSCAR-2301 is a multilingual web corpus extracted and filtered from the late-2022 Common Crawl snapshot, covering 151 languages and totaling ~3.57TB of raw unannotated text for LLM pretraining. It is published by the OSCAR project (oscar-corpus) and includes KenLM-based content quality metrics and deduplication hashes. The corpus content is from Common Crawl with no copyright held by the authors; metadata is CC0-1.0.
Verified live 2026-06-22 via primary sources. Manually gated and the card states access is temporarily suspended pending clarification.
Openness
3 high confidence- license
- cc0-1.0(metadata)
- gated
- manual(login + terms acceptance)
Manually gated and the card states access is temporarily suspended pending clarification.
- https://huggingface.co/datasets/oscar-corpus/OSCAR-2301 recorded 2026-06-22
Card describing 151 languages/~3.57TB, CC0 metadata, gated access with a notice suspending access
Adoption
2 high confidence1,793 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/oscar-corpus/OSCAR-2301 recorded 2026-07-04
downloads field = 1,793 (30-day)
Capability
3 high confidenceA dated 2022-snapshot web corpus with access suspended; the derived Colossal OSCAR trains ALIA/Salamandra (2024) but the 2301 version itself is superseded by FineWeb-2/CulturaX.
- https://huggingface.co/BSC-LT/salamandra-7b recorded 2026-07-04
Colossal OSCAR contributes 53% of Salamandra/ALIA pretraining tokens
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-07-04
FineWeb-2 supersedes raw OSCAR web extraction in quality ablations
Unchanged since 2026-07-04 (last edited, not re-checked)