AI Potluck
Product / UX / Safety & Guardrails

Aegis AI Content Safety

NVIDIA

NVIDIA's content-safety classifier (now also called Llama Nemotron Safety Guard Defensive), a LlamaGuard/Llama2-7B-based model parameter-efficiently tuned on NVIDIA's Aegis safety dataset to classify prompts and responses across 13 critical risk categories.

NVIDIA Aegis AI Content Safety (Defensive 1.0), aka Llama Nemotron Safety Guard Defensive V1. Llama 2 Community License (not OSI). Adapter weights public on HF, but the LlamaGuard base must be obtained via Meta's access request, so not fully open-weight. Aegis safety dataset documented. Verified June 2026.

Openness

3 medium confidence
3.0
weights
adapter-open(HF) but base LlamaGuard gated via Meta access
data
Aegis safety dataset documented
license
Llama-2-Community(NOT OSI)

Open-weights at the adapter level but the LlamaGuard base is access-gated through Meta and the license is the non-OSI Llama 2 Community License, so it lands at open_weights (3) with a gating caveat.

Adoption

2 low confidence
2.0

NVIDIA safety classifier with a documented dataset and NeMo integration; modest standalone HF adoption, distributed largely via NVIDIA's stack.

Capability

4 low confidence
4.0

Broad risk taxonomy with a documented training dataset; comparable coverage to Llama Guard, its base.

Unchanged since 2026-06-29 (last edited, not re-checked)