RedPajama-Data-v2
Together AIRedPajama-Data-v2 is a large open web dataset built from Common Crawl, containing over 100B text documents with around 30B documents annotated with quality signals and ~20B deduplicated documents (tens of trillions of tokens). It was created by Together Computer to support open language-model pretraining. The processing scripts are Apache-2.0, while the underlying text follows Common Crawl's Terms of Use.
Verified live 2026-06-22 via primary sources. Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.
Openness
5 high confidence- license
- CommonCrawl-ToU+Apache-2.0(code)
- ungated
- yes
- dataset_card
- yes
Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.
- https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2 recorded 2026-06-22
Common Crawl ToU + Apache-2.0 code license, ungated access, dataset card
Adoption
2 high confidence9,452 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/togethercomputer/RedPajama-Data-V2 recorded 2026-07-04
downloads field = 9,452 (30-day)
Capability
3 high confidenceCompetitive in its own filtered ablation but loses to FineWeb in FineWeb's 1.82B run; primarily a 2023 base (RedPajama-INCITE), superseded by DCLM/FineWeb.
- https://arxiv.org/abs/2411.12372 recorded 2026-07-04
RPv2 with Gopher filtering is competitive vs C4/Dolma but trails RefinedWeb
- https://arxiv.org/abs/2406.17557 recorded 2026-07-04
FineWeb outperforms RedPajama2 in the FineWeb ablation
Unchanged since 2026-07-04 (last edited, not re-checked)