AI Potluck
Product / UX / Safety & Guardrails

Zentropi CoPE

Zentropi

Zentropi's Content Policy Evaluator (CoPE-A, 9B, built on Gemma-2-9B with LoRA): a steerable, bring-your-own-policy content-labeling model. Rather than a fixed taxonomy, it evaluates content against natural-language policies you define, scoring across harm areas like hate speech, sexual content, violence, harassment, self-harm, and toxicity.

Zentropi CoPE-A (9B) on Gemma-2-9B (LoRA); license zentropi-openrail-m (OpenRAIL-M, use-restricted, NOT OSI). Open weights, self-hostable. RMC partner model distributed via the ROOST Model Community. 23 likes (June 2026); arXiv 2512.18027. Verified June 2026.

Openness

3 medium confidence
3.0
weights
open(zentropi-ai/cope-a-9b on HF)
base
Gemma-2-9B(LoRA)
license
zentropi-openrail-m(OpenRAIL-M, use-restricted, NOT OSI)

Open weights and self-hostable, but under an OpenRAIL-M license with use restrictions (e.g. surveillance), which is not OSI-approved, so open_weights (3).

Adoption

2 low confidence
2.0

Niche but notable: a ROOST Model Community partner model (distributed via RMC); 23 HF likes (June 2026), downloads not tracked.

Capability

4 low confidence
4.0

Like gpt-oss-safeguard, a bring-your-own-policy model that generalizes to custom policies; strong reported F1 across harm areas.

Unchanged since 2026-06-29 (last edited, not re-checked)