Llama Guard
MetaMeta's input/output safety classifier built on Llama 3.1-8B, the de-facto open guardrail model. It labels both prompts and responses across 14 hazard categories (violent crime, child exploitation, hate, self-harm, code-interpreter abuse, and more) and supports eight languages. Widely used as the moderation layer in front of open chat and agent stacks.
Llama Guard 3 (8B), built on Llama 3.1; Llama 3.1 Community License (not OSI; >700M MAU restriction). ~237K HF downloads/month, 307 likes (June 2026). Verified live June 2026.
Openness
3 medium confidence- weights
- open(Llama-Guard-3-8B safetensors on HF)
- data
- not-released
- code
- inference recipe public
- license
- Llama-3.1-Community(NOT OSI, >700M-MAU restriction)
Open-weights safety classifier under Meta's Llama Community License, which carries use restrictions and is not OSI-approved, so it lands at the open_weights tier (3) like the rest of the Llama family rather than open_source.
- https://huggingface.co/meta-llama/Llama-Guard-3-8B recorded 2026-06-29
Llama 3.1 Community License; open weights; 14 hazard categories; I/O classifier; ~237K downloads/month, 307 likes
Adoption
4 high confidence~237K Hugging Face downloads/month on the 8B variant (June 2026), the most-adopted open guardrail model; widely embedded as a moderation layer.
- https://huggingface.co/meta-llama/Llama-Guard-3-8B recorded 2026-06-29
~237,092 downloads/month, 307 likes (June 2026)
Capability
4 medium confidenceBroad hazard taxonomy and multilingual coverage; the reference open guardrail others are measured against.
- https://huggingface.co/meta-llama/Llama-Guard-3-8B recorded 2026-06-29
14 hazard categories, 8 languages, I/O moderation
Unchanged since 2026-06-29 (last edited, not re-checked)