Zentropi CoPE
ZentropiZentropi's Content Policy Evaluator (CoPE-A, 9B, built on Gemma-2-9B with LoRA): a steerable, bring-your-own-policy content-labeling model. Rather than a fixed taxonomy, it evaluates content against natural-language policies you define, scoring across harm areas like hate speech, sexual content, violence, harassment, self-harm, and toxicity.
Zentropi CoPE-A (9B) on Gemma-2-9B, trained as a LoRA adapter - the card has developers download the Gemma base and merge the adapter onto it, so the repo ships the adapter rather than a full model. An RMC partner model distributed through the ROOST Model Community; described in arXiv 2512.18027. Verified 2026-08-13 via the HF model card.
Openness
3 medium confidence- weights
- open(zentropi-ai/cope-a-9b on HF)
- base
- Gemma-2-9B(LoRA)
- license
- zentropi-openrail-m(OpenRAIL-M, use-restricted, NOT OSI)
Open weights and self-hostable, but under an OpenRAIL-M license with use restrictions (e.g. surveillance), which is not OSI-approved, so open_weights (3).
- https://huggingface.co/zentropi-ai/cope-a-9b recorded 2026-08-13
Repo header reads License - zentropi-openrail-m and cardData records license - other with license_name - zentropi-openrail-m. The repo is publicly accessible behind a contact-information gate rather than a private one, and ships adapter_model.safetensors; the card tells developers to download google/gemma2-9b and merge the CoPE adapter.
Adoption
2 low confidenceNiche but notable: a ROOST Model Community partner model, distributed through the RMC. The Hugging Face repo shows 23 likes and zero downloads over the trailing 30 days, because what it distributes is a LoRA adapter that most users merge into Gemma rather than pull directly - so there is no download figure to band on, and the level rests on reported traction instead. Zentropi publishes no usage count, so none is claimed here.
- https://huggingface.co/zentropi-ai/cope-a-9b recorded 2026-08-13
23 likes; downloads 0 over the trailing 30 days; no usage figure published anywhere on the card.
Capability
4 low confidenceLike gpt-oss-safeguard, a bring-your-own-policy model that evaluates content against arbitrary supplied policies rather than a fixed taxonomy, with strong reported F1 across the harm areas. The same shape of product as gpt-oss-safeguard, at the same rung.
- https://huggingface.co/zentropi-ai/cope-a-9b recorded 2026-08-13
Card lists 'Policy-adaptive content evaluation', 'Steerable (no fixed taxonomy/definitions)' and binary classification against a supplied policy, with training coverage across hate speech, sexual content, self-harm, harassment and toxicity; evaluation is largely on internal held-out policies besides the Ethos hate-speech benchmark.
Verified 2026-08-13