AI Potluck
Product / UX / Safety & Guardrails

gpt-oss-safeguard

OpenAI

OpenAI's open safety reasoning model, built on gpt-oss and released with Hugging Face and ROOST. Rather than emitting a fixed label, it reasons over a safety policy you supply and returns a reasoned decision, with configurable reasoning effort. Shipped in 20b and 120b variants under Apache 2.0.

gpt-oss-safeguard (20b / 120b), Apache-2.0 open weights; released by OpenAI with Hugging Face and ROOST, distributed via the ROOST Model Community (RMC). 20b ~131K HF downloads/month, 235 likes (June 2026). Verified live June 2026.

Openness

4 medium confidence
4.0
weights
open(gpt-oss-safeguard-20b/120b on HF)
data
not-released
license
Apache-2.0(OSI)

Open weights under Apache-2.0 (OSI), training data not released, so open_weights (4) rather than open_source. Notable as a policy-conditioned safety reasoning model rather than a fixed classifier.

Adoption

4 high confidence
4.0

20b variant ~131K HF downloads/month, 235 likes, 100+ Spaces (June 2026); strong early adoption for a recent release.

Capability

4 medium confidence
4.0

A different shape from fixed classifiers: reasons over an arbitrary supplied policy, which generalizes to novel categories at some latency cost.

Unchanged since 2026-06-29 (last edited, not re-checked)