gpt-oss-safeguard
OpenAIOpenAI's open safety reasoning model, built on gpt-oss and released with Hugging Face and ROOST. Rather than emitting a fixed label, it reasons over a safety policy you supply and returns a reasoned decision, with configurable reasoning effort. Shipped in 20b and 120b variants under Apache 2.0.
gpt-oss-safeguard (20b / 120b), Apache-2.0 open weights; released by OpenAI with Hugging Face and ROOST, distributed via the ROOST Model Community (RMC). 20b ~131K HF downloads/month, 235 likes (June 2026). Verified live June 2026.
Openness
4 medium confidence- weights
- open(gpt-oss-safeguard-20b/120b on HF)
- data
- not-released
- license
- Apache-2.0(OSI)
Open weights under Apache-2.0 (OSI), training data not released, so open_weights (4) rather than open_source. Notable as a policy-conditioned safety reasoning model rather than a fixed classifier.
- https://huggingface.co/openai/gpt-oss-safeguard-20b recorded 2026-06-29
Apache-2.0; open weights; policy-conditioned safety reasoning; 20b/120b; ~131K downloads/month, 235 likes
Adoption
4 high confidence20b variant ~131K HF downloads/month, 235 likes, 100+ Spaces (June 2026); strong early adoption for a recent release.
- https://huggingface.co/openai/gpt-oss-safeguard-20b recorded 2026-06-29
~131,038 downloads/month, 235 likes, 100+ Spaces (June 2026)
Capability
4 medium confidenceA different shape from fixed classifiers: reasons over an arbitrary supplied policy, which generalizes to novel categories at some latency cost.
- https://huggingface.co/openai/gpt-oss-safeguard-20b recorded 2026-06-29
reasoned decisions, configurable reasoning effort, bring-your-own-policy
Unchanged since 2026-06-29 (last edited, not re-checked)