OSCAR-2301
OSCAR ProjectMultilingual web corpus extracted and filtered from the late-2022 Common Crawl snapshot, covering 151 languages and totalling about 3.57TB of raw unannotated text for pretraining. Published by the OSCAR project, it includes KenLM-based content quality metrics and deduplication hashes.
The card states that access is temporarily suspended pending clarification, so the data itself could not be reached for this record. Verified 2026-08-13 via the OSCAR-2301 dataset card on Hugging Face.
Openness
3 high confidence- license
- cc0-1.0(metadata)
- gated
- manual(login + terms acceptance)
Manually gated and the card states access is temporarily suspended pending clarification.
- https://huggingface.co/datasets/oscar-corpus/OSCAR-2301 recorded 2026-08-13
Card describing 151 languages/~3.57TB, CC0 metadata, gated access with a notice suspending access
Adoption
2 high confidence2,036 downloads in the trailing 30 days for oscar-corpus/OSCAR-2301, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/oscar-corpus/OSCAR-2301 recorded 2026-08-13
2,036 downloads in the trailing 30 days for oscar-corpus/OSCAR-2301
Capability
3 high confidenceA dated 2022-snapshot web corpus with access suspended; the derived Colossal OSCAR trains ALIA/Salamandra (2024) but the 2301 version itself is superseded by FineWeb-2/CulturaX. The Salamandra card puts Colossal OSCAR at 53.05% of its pretraining tokens.
- https://huggingface.co/BSC-LT/salamandra-7b recorded 2026-08-13
Colossal OSCAR contributes 53% of Salamandra/ALIA pretraining tokens
- https://huggingface.co/datasets/HuggingFaceFW/fineweb-2 recorded 2026-08-13
FineWeb-2 supersedes raw OSCAR web extraction in quality ablations
Verified 2026-08-13