AI Potluck
Product / UX / Safety & Guardrails

Aegis AI Content Safety

NVIDIA

NVIDIA's content-safety classifier (now also called Llama Nemotron Safety Guard Defensive), a LlamaGuard/Llama2-7B-based model parameter-efficiently tuned on NVIDIA's Aegis safety dataset to classify prompts and responses across 13 critical risk categories.

NVIDIA Aegis AI Content Safety (Defensive 1.0), also called Llama Nemotron Safety Guard Defensive V1. The adapter weights are on the Hub while the Llama Guard base needs Meta's access request, which the openness axis weighs. The Aegis dataset it was tuned on is published. Verified 2026-08-13 via the HF model card.

Openness

3 medium confidence
3.0
weights
adapter-open(HF) but base LlamaGuard gated via Meta access
data
Aegis safety dataset documented
license
Llama-2-Community(NOT OSI)

Open-weights at the adapter level but the LlamaGuard base is access-gated through Meta and the license is the non-OSI Llama 2 Community License, so it lands at open_weights (3) with a gating caveat.

Adoption

1 low confidence
1.0

4,303 downloads in the trailing 30 days for the one declared artifact, nvidia/Aegis-AI-Content-Safety-LlamaGuard-Defensive-1.0, which bands at under 10K, level 1 on the software and model adoption scale.

Capability

4 low confidence
4.0

Broad risk taxonomy with a documented training dataset. Aegis is itself a tune of Llama Guard and covers a taxonomy of the same order - 13 categories against Llama Guard's 14 - so it sits at the same rung as its base.

Verified 2026-08-13