AI Potluck
Product / UX / Safety & Guardrails

Llama Guard

Meta

Meta's input/output safety classifier built on Llama 3.1-8B, the de-facto open guardrail model. It labels both prompts and responses across 14 hazard categories (violent crime, child exploitation, hate, self-harm, code-interpreter abuse, and more) and supports eight languages. Widely used as the moderation layer in front of open chat and agent stacks.

Llama Guard 3 (8B), built on Llama 3.1; Llama 3.1 Community License (not OSI; >700M MAU restriction). ~237K HF downloads/month, 307 likes (June 2026). Verified live June 2026.

Openness

3 medium confidence
3.0
weights
open(Llama-Guard-3-8B safetensors on HF)
data
not-released
code
inference recipe public
license
Llama-3.1-Community(NOT OSI, >700M-MAU restriction)

Open-weights safety classifier under Meta's Llama Community License, which carries use restrictions and is not OSI-approved, so it lands at the open_weights tier (3) like the rest of the Llama family rather than open_source.

Adoption

4 high confidence
4.0

~237K Hugging Face downloads/month on the 8B variant (June 2026), the most-adopted open guardrail model; widely embedded as a moderation layer.

Capability

4 medium confidence
4.0

Broad hazard taxonomy and multilingual coverage; the reference open guardrail others are measured against.

Unchanged since 2026-06-29 (last edited, not re-checked)