AI Potluck
Product / UX / Safety & Guardrails

Granite Guardian

IBM

IBM's open guardrail model family for detecting risks in prompts and responses, plus RAG-specific checks. It covers harm, social bias, jailbreaking, violence, profanity and sexual content, and adds hallucination checks for retrieval pipelines (context relevance, groundedness, answer relevance).

Granite Guardian 3.2 (5B, distilled from the 8B), covering harm, social bias, jailbreaking, violence, profanity and sexual content, plus RAG context relevance, groundedness and answer relevance and function-calling risk. The card describes the training data as human-annotated and synthetic without publishing the mixture. Verified 2026-08-13 via the HF model card.

Openness

3 medium confidence
3.0
weights
open(granite-guardian-3.2-5b on HF)
data
not-released
license
Apache-2.0(OSI)

Open weights under Apache-2.0, permissive and OSI-approved, but neither the post-training data nor the fine-tuning pipeline is released, and the rung above requires both - so this scores 3, the rung every comparable Apache-2.0 guardrail model sits on. The card has a Training Data section that names the ingredients without publishing the mixture. A strong open-weights release even so, and the licence is not what holds it back.

Adoption

1 medium confidence
1.0

1,837 downloads in the trailing 30 days for the one declared artifact, ibm-granite/granite-guardian-3.2-5b, which bands at under 10K, level 1 on the software and model adoption scale.

Capability

4 medium confidence
4.0

One of the broadest open guardrails: standard harm taxonomy plus RAG-specific groundedness/hallucination checks.

  • https://huggingface.co/ibm-granite/granite-guardian-3.2-5b recorded 2026-08-13

    Card enumerates Harm, Social Bias, Jailbreaking, Violence, Profanity and Sexual Content, plus RAG hallucination risks - context relevance, groundedness and answer relevance - and function-calling risk detection in agentic workflows.

Verified 2026-08-13