Llama Prompt Guard 2
MetaMeta's compact prompt-injection and jailbreak classifier (86M and 22M), labeling prompts benign or malicious. Part of LlamaFirewall, it is tuned for very low latency (sub-100ms) and high precision, with multilingual coverage across eight languages.
Llama Prompt Guard 2 (86M / 22M); Llama 4 Community License (not OSI). 99.8% AUC, 97.5% recall at 1% FPR on English jailbreaks; 8 languages. 86M ~90.4K HF downloads/month, 149 likes (June 2026). Verified live June 2026.
Openness
3 medium confidence- weights
- open(Llama-Prompt-Guard-2-86M/22M on HF)
- data
- not-released
- license
- Llama-4-Community(NOT OSI)
Open weights under Meta's Llama 4 Community License, which is use-restricted and not OSI-approved, so open_weights (3) like the rest of the Llama family.
- https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M recorded 2026-06-29
Llama 4 Community License; open weights; 86M/22M; prompt-injection/jailbreak classifier; ~90,441 downloads/month, 149 likes
Adoption
4 high confidence86M variant ~90.4K HF downloads/month (June 2026); widely used as a low-latency injection filter, including inside LlamaFirewall.
- https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M recorded 2026-06-29
~90,441 downloads/month, 149 likes (June 2026)
Capability
4 medium confidenceSpecialized injection/jailbreak classifier with strong accuracy and very low latency; complements full-taxonomy guards.
- https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M recorded 2026-06-29
99.8% AUC, 97.5% recall @1% FPR; 8 languages; sub-100ms
Unchanged since 2026-06-29 (last edited, not re-checked)