AI Potluck
Model components / Training & synthetic datasets

Nemotron-CC

NVIDIA

NVIDIA's Common Crawl-derived pretraining corpus family, spanning trillions of tokens with synthetic rephrasing across general, math, code and specialized sets. It backs the Nemotron model line and is a documented component of Marin 8B.

A family record covering the v1 and v2 corpora and the math, code and specialized sets. Verified 2026-08-13 via the nvidia/Nemotron-CC-v2 dataset card on Hugging Face.

Openness

3 high confidence
3.0
license
NVIDIA-Open-Data
gated
manual(model-training-use)
dataset_card
present

Manually gated behind the NVIDIA Data Agreement for Model Training; synthetic portions may inherit generator-model license terms.

Adoption

3 high confidence
3.0

45,663 downloads in the trailing 30 days for nvidia/Nemotron-CC-v2, which bands at 10K-100K, level 3 on the dataset adoption scale.

Capability

5 high confidence
5.0

Wins against DCLM at a long token horizon: Nemotron-CC-HQ adds 5.6 MMLU over DCLM, and an 8B trained on 15T tokens of it beats Llama 3.1 8B by 5 MMLU. It is the current default for NVIDIA's Nemotron models, and the Marin retrospective records Marin 8B adopting it as its primary source in later phases.

Verified 2026-08-13