RedPajama-Data-v2
Together AILarge open web dataset built from Common Crawl by Together Computer, holding more than 100 billion text documents, of which around 30 billion carry quality-signal annotations and about 20 billion are deduplicated. It is published raw and unfiltered so downstream users apply their own filtering.
Verified 2026-08-13 via the togethercomputer/RedPajama-Data-V2 dataset card on Hugging Face.
Openness
5 high confidence- license
- CommonCrawl-ToU+Apache-2.0(code)
- ungated
- yes
- dataset_card
- yes
Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.
- https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2 recorded 2026-08-13
Common Crawl ToU + Apache-2.0 code license, ungated access, dataset card
Adoption
2 high confidence5,035 downloads in the trailing 30 days for togethercomputer/RedPajama-Data-V2, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/togethercomputer/RedPajama-Data-V2 recorded 2026-08-13
5,035 downloads in the trailing 30 days for togethercomputer/RedPajama-Data-V2
Capability
3 high confidenceCompetitive in its own filtered ablation but loses to FineWeb in FineWeb's 1.82B run; primarily a 2023 base (RedPajama-INCITE), superseded by DCLM/FineWeb. V2 ships as raw unfiltered web text plus quality signals rather than as a filtered corpus.
- https://arxiv.org/abs/2411.12372 recorded 2026-08-13
Abstract - RedPajama-V2 is a raw, unfiltered web corpus of over 100T tokens shipped with quality signals for downstream filtering. The Gopher-filtered ablation against C4/Dolma/RefinedWeb is in the paper body, not on this page
- https://arxiv.org/abs/2406.17557 recorded 2026-08-13
FineWeb outperforms RedPajama2 in the FineWeb ablation
Verified 2026-08-13