Granite Guardian
IBMIBM's open guardrail model family for detecting risks in prompts and responses, plus RAG-specific checks. It covers harm, social bias, jailbreaking, violence, profanity and sexual content, and adds hallucination checks for retrieval pipelines (context relevance, groundedness, answer relevance).
Granite Guardian 3.2 (5B, distilled from the 8B); Apache-2.0 open weights. ~1.4K HF downloads/month, 14 likes (June 2026). Verified live June 2026.
Openness
4 medium confidence- weights
- open(granite-guardian-3.2-5b on HF)
- data
- not-released
- license
- Apache-2.0(OSI)
Open weights under Apache-2.0 (permissive, OSI), but the training data is not released, so it stops short of the open_source (5) tier; a strong open_weights release at 4.
- https://huggingface.co/ibm-granite/granite-guardian-3.2-5b recorded 2026-06-29
Apache-2.0; open weights; harm/bias/jailbreak + RAG groundedness checks; ~1,377 downloads/month, 14 likes
Adoption
2 medium confidence~1.4K HF downloads/month on the 3.2-5B variant (June 2026); cumulative across the Guardian family is higher. Enterprise-leaning distribution via watsonx.
- https://huggingface.co/ibm-granite/granite-guardian-3.2-5b recorded 2026-06-29
~1,377 downloads/month, 14 likes (June 2026)
Capability
4 medium confidenceOne of the broadest open guardrails: standard harm taxonomy plus RAG-specific groundedness/hallucination checks.
- https://huggingface.co/ibm-granite/granite-guardian-3.2-5b recorded 2026-06-29
harm/bias/jailbreak + RAG groundedness/answer-relevance checks
Unchanged since 2026-06-29 (last edited, not re-checked)