AI Potluck
Model components / Training & synthetic datasets

OSCAR-2301

OSCAR Project

Multilingual web corpus extracted and filtered from the late-2022 Common Crawl snapshot, covering 151 languages and totalling about 3.57TB of raw unannotated text for pretraining. Published by the OSCAR project, it includes KenLM-based content quality metrics and deduplication hashes.

The card states that access is temporarily suspended pending clarification, so the data itself could not be reached for this record. Verified 2026-08-13 via the OSCAR-2301 dataset card on Hugging Face.

Openness

3 high confidence
3.0
license
cc0-1.0(metadata)
gated
manual(login + terms acceptance)

Manually gated and the card states access is temporarily suspended pending clarification.

Adoption

2 high confidence
2.0

2,036 downloads in the trailing 30 days for oscar-corpus/OSCAR-2301, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

3 high confidence
3.0

A dated 2022-snapshot web corpus with access suspended; the derived Colossal OSCAR trains ALIA/Salamandra (2024) but the 2301 version itself is superseded by FineWeb-2/CulturaX. The Salamandra card puts Colossal OSCAR at 53.05% of its pretraining tokens.

Verified 2026-08-13