AI Potluck
Model components / Training & synthetic datasets

RedPajama-Data-v2

Together AI

RedPajama-Data-v2 is a large open web dataset built from Common Crawl, containing over 100B text documents with around 30B documents annotated with quality signals and ~20B deduplicated documents (tens of trillions of tokens). It was created by Together Computer to support open language-model pretraining. The processing scripts are Apache-2.0, while the underlying text follows Common Crawl's Terms of Use.

Verified live 2026-06-22 via primary sources. Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.

Openness

5 high confidence
5.0
license
CommonCrawl-ToU+Apache-2.0(code)
ungated
yes
dataset_card
yes

Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.

Adoption

2 high confidence
2.0

9,452 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

3 high confidence
3.0

Competitive in its own filtered ablation but loses to FineWeb in FineWeb's 1.82B run; primarily a 2023 base (RedPajama-INCITE), superseded by DCLM/FineWeb.

Unchanged since 2026-07-04 (last edited, not re-checked)