Nemotron-CC
NVIDIANVIDIA's Common Crawl-derived pretraining corpus family, spanning trillions of tokens with synthetic rephrasing across general, math, code and specialized sets. It backs the Nemotron model line and is a documented component of Marin 8B.
A family record covering the v1 and v2 corpora and the math, code and specialized sets. Verified 2026-08-13 via the nvidia/Nemotron-CC-v2 dataset card on Hugging Face.
Openness
3 high confidence- license
- NVIDIA-Open-Data
- gated
- manual(model-training-use)
- dataset_card
- present
Manually gated behind the NVIDIA Data Agreement for Model Training; synthetic portions may inherit generator-model license terms.
- https://huggingface.co/datasets/nvidia/Nemotron-CC-v2 recorded 2026-08-13
gated (manual), NVIDIA Open Data License, dataset card
Adoption
3 high confidence45,663 downloads in the trailing 30 days for nvidia/Nemotron-CC-v2, which bands at 10K-100K, level 3 on the dataset adoption scale.
- https://huggingface.co/api/datasets/nvidia/Nemotron-CC-v2 recorded 2026-08-13
45,663 downloads in the trailing 30 days for nvidia/Nemotron-CC-v2
Capability
5 high confidenceWins against DCLM at a long token horizon: Nemotron-CC-HQ adds 5.6 MMLU over DCLM, and an 8B trained on 15T tokens of it beats Llama 3.1 8B by 5 MMLU. It is the current default for NVIDIA's Nemotron models, and the Marin retrospective records Marin 8B adopting it as its primary source in later phases.
- https://arxiv.org/abs/2412.02595 recorded 2026-08-13
Nemotron-CC-HQ improves MMLU by 5.6 over DCLM; an 8B on 15T tokens beats Llama 3.1 8B by +5 MMLU
- https://marin.readthedocs.io/en/latest/reports/marin-8b-retro/ recorded 2026-08-13
Marin 8B adopted Nemotron-CC as its primary source in later phases
Verified 2026-08-13