AI Potluck
Model components / Dataset Processing Tools

NeMo Curator

NVIDIA

GPU-accelerated, scalable data-curation toolkit for LLM training corpora across text, image, video, and audio: extraction, language ID, exact/fuzzy/substring dedup, quality classification, and synthetic generation.

NVIDIA NeMo Curator; repo actively maintained (v1.2.0, 2026-05-14), Apache-2.0 confirmed live 2026-07-21. Powers the Nemotron-CC curation pipeline and Nemotron-4 data prep (8T+ multilingual tokens).

Openness

5 high confidence
5.0
license
Apache-2.0(OSI)
source
public(NVIDIA/NeMo-Curator)
governance
NVIDIA
core-gated
ungated

Fully OSI-licensed (Apache-2.0), full source public. Vendor-backed but not gated.

Adoption

2 high confidence
2.0

~5.6k PyPI downloads/month; NVIDIA-backed and the pipeline behind the Nemotron-CC family, but modest raw download volume.

Capability

5 high confidence
5.0

Benchmark-anchored on the output corpus (proxy): Nemotron-CC matches or beats the DCLM/FineWeb-Edu frontier, especially at long-horizon token budgets.

Unchanged since 2026-07-30 (last edited, not re-checked)