AI Potluck
Model components / Dataset Processing Tools

NeMo Curator

NVIDIA

GPU-accelerated, scalable data-curation toolkit for LLM training corpora across text, image, video, and audio: extraction, language ID, exact/fuzzy/substring dedup, quality classification, and synthetic generation.

Powers the Nemotron-CC curation pipeline and Nemotron-4 data prep (8T+ multilingual tokens). The repo now lives at NVIDIA-NeMo/Curator and the older NVIDIA/NeMo-Curator path redirects to it. Verified 2026-08-13 via the GitHub README and the LICENSE body.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(NVIDIA/NeMo-Curator)
governance
NVIDIA
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public. Vendor-backed but not gated.

Adoption

1 high confidence
1.0

4,075 downloads in the trailing 30 days for the declared `nemo-curator` package, which bands at <10K, level 1 on the software and model adoption scale.

Capability

5 high confidence
5.0

Benchmark-anchored on the output corpus (proxy): Nemotron-CC matches or beats the DCLM/FineWeb-Edu frontier, especially at long-horizon token budgets.

  • https://arxiv.org/abs/2412.02595 recorded 2026-08-13

    Nemotron-CC abstract - a high-quality subset improves MMLU by 5.6 over DCLM at 8B/1T, the full 6.3T-token set matches DCLM on MMLU with four times more unique real tokens, and an 8B model trained on 15T tokens beats Llama 3.1 8B by +5 MMLU

Verified 2026-08-13