LlamaFirewall
MetaMeta's open, system-level guardrail framework for AI agents (part of PurpleLlama). It composes layered defenses: Prompt Guard 2 (jailbreak/injection detection), Agent Alignment Checks (detecting goal hijacking), and CodeShield (filtering insecure generated code), into a real-time firewall around agent execution.
LlamaFirewall, in meta-llama/PurpleLlama (~4.2K stars). PurpleLlama licensing is mixed: evals/benchmarks MIT, safeguard models under the Llama Community License. Components: Prompt Guard 2, Agent Alignment Checks, CodeShield. Verified live June 2026.
Openness
5 high confidence- framework
- open(LlamaFirewall + CodeShield code, MIT-licensed in PurpleLlama)
- self-host
- yes
- note
- bundled guard models (Prompt Guard 2, Llama Guard) carry the Llama Community License
The firewall framework and CodeShield are open-source (MIT) and self-hostable; it orchestrates Meta's Llama-licensed guard models. Scored on the framework, which is genuinely open source (5).
- https://github.com/meta-llama/PurpleLlama recorded 2026-06-29
PurpleLlama: MIT evals + Llama-licensed safeguard models; ~4.2k stars; LlamaFirewall, Prompt Guard, CodeShield components
Adoption
3 low confidencePurpleLlama ~4.2K GitHub stars (June 2026); stars-only proxy, capped per methodology.
- https://github.com/meta-llama/PurpleLlama recorded 2026-06-29
~4.2k stars (June 2026)
Capability
4 medium confidenceOne of the few guardrail systems aimed specifically at agent security, with alignment checks and code filtering beyond I/O moderation.
- https://github.com/meta-llama/PurpleLlama recorded 2026-06-29
Prompt Guard 2 + Agent Alignment Checks + CodeShield
Unchanged since 2026-06-29 (last edited, not re-checked)