AI Potluck
Model components / Training & synthetic datasets

Nemotron-CC

NVIDIA

Nemotron-CC is NVIDIA's Common Crawl-derived pretraining corpus family (v1/v2 plus Math, Code, and Specialized sets), spanning trillions of tokens with synthetic rephrasing, released under the NVIDIA Open Data License with gated access. It backs the Nemotron model line and is a documented component of Marin 8B.

Family bucket: Nemotron-CC v1/v2, Nemotron-CC-Math, Nemotron-Pretraining-Code, and Nemotron-Pretraining-Specialized.

Openness

3 high confidence
3.0
license
NVIDIA-Open-Data
gated
manual(model-training-use)
dataset_card
present

Manually gated behind the NVIDIA Data Agreement for Model Training; synthetic portions may inherit generator-model license terms.

Adoption

3 high confidence
3.0

16,100 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

5 high confidence
5.0

Wins vs DCLM at a long token horizon (Nemotron-CC-HQ +5.6 MMLU over DCLM; a 15T-token 8B beats Llama 3.1 8B); current default for NVIDIA Nemotron models and Marin 8B late phases.

Unchanged since 2026-07-04 (last edited, not re-checked)