NeMo Curator
NVIDIAGPU-accelerated, scalable data-curation toolkit for LLM training corpora across text, image, video, and audio: extraction, language ID, exact/fuzzy/substring dedup, quality classification, and synthetic generation.
NVIDIA NeMo Curator; repo actively maintained (v1.2.0, 2026-05-14), Apache-2.0 confirmed live 2026-07-21. Powers the Nemotron-CC curation pipeline and Nemotron-4 data prep (8T+ multilingual tokens).
Openness
5 high confidence- license
- Apache-2.0(OSI)
- source
- public(NVIDIA/NeMo-Curator)
- governance
- NVIDIA
- core-gated
- ungated
Fully OSI-licensed (Apache-2.0), full source public. Vendor-backed but not gated.
- https://github.com/NVIDIA/NeMo-Curator/blob/main/LICENSE recorded 2026-07-21
Apache License Version 2.0 text
Adoption
2 high confidence~5.6k PyPI downloads/month; NVIDIA-backed and the pipeline behind the Nemotron-CC family, but modest raw download volume.
- https://pypistats.org/packages/nemo-curator recorded 2026-07-21
Downloads last month: 5,618
Capability
5 high confidenceBenchmark-anchored on the output corpus (proxy): Nemotron-CC matches or beats the DCLM/FineWeb-Edu frontier, especially at long-horizon token budgets.
- https://arxiv.org/abs/2412.02595 recorded 2026-07-21
Nemotron-CC-HQ +5.6 MMLU over DCLM at 8B/1T; 70.3 vs 65.3 MMLU at 15T tokens
Unchanged since 2026-07-30 (last edited, not re-checked)