Aegis AI Content Safety
NVIDIANVIDIA's content-safety classifier (now also called Llama Nemotron Safety Guard Defensive), a LlamaGuard/Llama2-7B-based model parameter-efficiently tuned on NVIDIA's Aegis safety dataset to classify prompts and responses across 13 critical risk categories.
NVIDIA Aegis AI Content Safety (Defensive 1.0), aka Llama Nemotron Safety Guard Defensive V1. Llama 2 Community License (not OSI). Adapter weights public on HF, but the LlamaGuard base must be obtained via Meta's access request, so not fully open-weight. Aegis safety dataset documented. Verified June 2026.
Openness
3 medium confidence- weights
- adapter-open(HF) but base LlamaGuard gated via Meta access
- data
- Aegis safety dataset documented
- license
- Llama-2-Community(NOT OSI)
Open-weights at the adapter level but the LlamaGuard base is access-gated through Meta and the license is the non-OSI Llama 2 Community License, so it lands at open_weights (3) with a gating caveat.
- https://huggingface.co/nvidia/Aegis-AI-Content-Safety-LlamaGuard-Defensive-1.0 recorded 2026-06-29
Llama 2 Community License; adapter weights public, base gated via Meta; 13 risk categories; built on LlamaGuard/Llama2-7B
Adoption
2 low confidenceNVIDIA safety classifier with a documented dataset and NeMo integration; modest standalone HF adoption, distributed largely via NVIDIA's stack.
- https://huggingface.co/nvidia/Aegis-AI-Content-Safety-LlamaGuard-Defensive-1.0 recorded 2026-06-29
NVIDIA Aegis content-safety classifier; 13 categories
Capability
4 low confidenceBroad risk taxonomy with a documented training dataset; comparable coverage to Llama Guard, its base.
- https://huggingface.co/nvidia/Aegis-AI-Content-Safety-LlamaGuard-Defensive-1.0 recorded 2026-06-29
13 risk categories; defensive variant
Unchanged since 2026-06-29 (last edited, not re-checked)