Nemotron-CC
NVIDIANemotron-CC is NVIDIA's Common Crawl-derived pretraining corpus family (v1/v2 plus Math, Code, and Specialized sets), spanning trillions of tokens with synthetic rephrasing, released under the NVIDIA Open Data License with gated access. It backs the Nemotron model line and is a documented component of Marin 8B.
Family bucket: Nemotron-CC v1/v2, Nemotron-CC-Math, Nemotron-Pretraining-Code, and Nemotron-Pretraining-Specialized.
Openness
3 high confidence- license
- NVIDIA-Open-Data
- gated
- manual(model-training-use)
- dataset_card
- present
Manually gated behind the NVIDIA Data Agreement for Model Training; synthetic portions may inherit generator-model license terms.
- https://huggingface.co/datasets/nvidia/Nemotron-CC-v2 recorded 2026-07-04
gated (manual), NVIDIA Open Data License, dataset card
Adoption
3 high confidence16,100 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/nvidia/Nemotron-CC-v2 recorded 2026-07-04
downloads field = 16,100 (30-day)
Capability
5 high confidenceWins vs DCLM at a long token horizon (Nemotron-CC-HQ +5.6 MMLU over DCLM; a 15T-token 8B beats Llama 3.1 8B); current default for NVIDIA Nemotron models and Marin 8B late phases.
- https://arxiv.org/abs/2412.02595 recorded 2026-07-04
Nemotron-CC-HQ improves MMLU by 5.6 over DCLM; an 8B on 15T tokens beats Llama 3.1 8B by +5 MMLU
- https://marin.readthedocs.io/en/latest/reports/marin-8b-retro/ recorded 2026-07-04
Marin 8B adopted Nemotron-CC as its primary source in later phases
Unchanged since 2026-07-04 (last edited, not re-checked)