AI Potluck
Product / UX / Safety & Guardrails

Llama Prompt Guard 2

Meta

Meta's compact prompt-injection and jailbreak classifier (86M and 22M), labeling prompts benign or malicious. Part of LlamaFirewall, it is tuned for very low latency (sub-100ms) and high precision, with multilingual coverage across eight languages.

Llama Prompt Guard 2 (86M / 22M); Llama 4 Community License (not OSI). 99.8% AUC, 97.5% recall at 1% FPR on English jailbreaks; 8 languages. 86M ~90.4K HF downloads/month, 149 likes (June 2026). Verified live June 2026.

Openness

3 medium confidence
3.0
weights
open(Llama-Prompt-Guard-2-86M/22M on HF)
data
not-released
license
Llama-4-Community(NOT OSI)

Open weights under Meta's Llama 4 Community License, which is use-restricted and not OSI-approved, so open_weights (3) like the rest of the Llama family.

Adoption

4 high confidence
4.0

86M variant ~90.4K HF downloads/month (June 2026); widely used as a low-latency injection filter, including inside LlamaFirewall.

Capability

4 medium confidence
4.0

Specialized injection/jailbreak classifier with strong accuracy and very low latency; complements full-taxonomy guards.

Unchanged since 2026-06-29 (last edited, not re-checked)