NeMo Curator
NVIDIAGPU-accelerated, scalable data-curation toolkit for LLM training corpora across text, image, video, and audio: extraction, language ID, exact/fuzzy/substring dedup, quality classification, and synthetic generation.
Powers the Nemotron-CC curation pipeline and Nemotron-4 data prep (8T+ multilingual tokens). The repo now lives at NVIDIA-NeMo/Curator and the older NVIDIA/NeMo-Curator path redirects to it. Verified 2026-08-13 via the GitHub README and the LICENSE body.
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(NVIDIA/NeMo-Curator)
- governance
- NVIDIA
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public. Vendor-backed but not gated.
- https://github.com/NVIDIA/NeMo-Curator/blob/main/LICENSE recorded 2026-08-13
Apache License Version 2.0 text
- https://raw.githubusercontent.com/NVIDIA-NeMo/Curator/main/README.md recorded 2026-08-13
README documents the whole pipeline building and installing from the published repo; no paid tier, enterprise edition or license key appears anywhere in it, so nothing is withheld from the source
Adoption
1 high confidence4,075 downloads in the trailing 30 days for the declared `nemo-curator` package, which bands at <10K, level 1 on the software and model adoption scale.
- https://pypistats.org/api/packages/nemo-curator/recent recorded 2026-08-13
4,075 downloads in the trailing 30 days for nemo-curator
Capability
5 high confidenceBenchmark-anchored on the output corpus (proxy): Nemotron-CC matches or beats the DCLM/FineWeb-Edu frontier, especially at long-horizon token budgets.
- https://arxiv.org/abs/2412.02595 recorded 2026-08-13
Nemotron-CC abstract - a high-quality subset improves MMLU by 5.6 over DCLM at 8B/1T, the full 6.3T-token set matches DCLM on MMLU with four times more unique real tokens, and an 8B model trained on 15T tokens beats Llama 3.1 8B by +5 MMLU
Verified 2026-08-13