AI Potluck
Product / UX / Safety & Guardrails

LlamaFirewall

Meta

Meta's open, system-level guardrail framework for AI agents (part of PurpleLlama). It composes layered defenses: Prompt Guard 2 (jailbreak/injection detection), Agent Alignment Checks (detecting goal hijacking), and CodeShield (filtering insecure generated code), into a real-time firewall around agent execution.

LlamaFirewall, in meta-llama/PurpleLlama (~4.2K stars). PurpleLlama licensing is mixed: evals/benchmarks MIT, safeguard models under the Llama Community License. Components: Prompt Guard 2, Agent Alignment Checks, CodeShield. Verified live June 2026.

Openness

5 high confidence
5.0
framework
open(LlamaFirewall + CodeShield code, MIT-licensed in PurpleLlama)
self-host
yes
note
bundled guard models (Prompt Guard 2, Llama Guard) carry the Llama Community License

The firewall framework and CodeShield are open-source (MIT) and self-hostable; it orchestrates Meta's Llama-licensed guard models. Scored on the framework, which is genuinely open source (5).

Adoption

3 low confidence
3.0

PurpleLlama ~4.2K GitHub stars (June 2026); stars-only proxy, capped per methodology.

Capability

4 medium confidence
4.0

One of the few guardrail systems aimed specifically at agent security, with alignment checks and code filtering beyond I/O moderation.

Unchanged since 2026-06-29 (last edited, not re-checked)