AI Potluck
Model components / Training & synthetic datasets

RedPajama-Data-v2

Together AI

Large open web dataset built from Common Crawl by Together Computer, holding more than 100 billion text documents, of which around 30 billion carry quality-signal annotations and about 20 billion are deduplicated. It is published raw and unfiltered so downstream users apply their own filtering.

Verified 2026-08-13 via the togethercomputer/RedPajama-Data-V2 dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
CommonCrawl-ToU+Apache-2.0(code)
ungated
yes
dataset_card
yes

Ungated with a full dataset card; redistributable per Common Crawl ToU with Apache-2.0 processing code.

Adoption

2 high confidence
2.0

5,035 downloads in the trailing 30 days for togethercomputer/RedPajama-Data-V2, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

3 high confidence
3.0

Competitive in its own filtered ablation but loses to FineWeb in FineWeb's 1.82B run; primarily a 2023 base (RedPajama-INCITE), superseded by DCLM/FineWeb. V2 ships as raw unfiltered web text plus quality signals rather than as a filtered corpus.

  • https://arxiv.org/abs/2411.12372 recorded 2026-08-13

    Abstract - RedPajama-V2 is a raw, unfiltered web corpus of over 100T tokens shipped with quality signals for downstream filtering. The Gopher-filtered ablation against C4/Dolma/RefinedWeb is in the paper body, not on this page

  • https://arxiv.org/abs/2406.17557 recorded 2026-08-13

    FineWeb outperforms RedPajama2 in the FineWeb ablation

Verified 2026-08-13