AI Potluck
Model components / Evaluation code

Patronus Evaluation Platform

Patronus AI

Managed evaluation platform built around proprietary judge models, Lynx (8B hallucination detector), GLIDER (3.8B explainable rubric judge) and Percival (agent execution-trace analyzer). Picked by enterprises needing managed eval models rather than spinning up their own LLM-as-judge. Raised $20M; usage-based pricing $10-$20 per 1,000 evaluator API calls.

Patronus AI hosted GenAI evaluation/optimization platform ('Core Platform'); open source client SDKs (patronus-py, latest v0.1.25 Jan 2026; patronus-api-node) that connect to proprietary hosted infra. Built-in evaluators Lynx (hallucination), Glider; experiments, production monitoring/tracing, Percival, RL Envs. Confirmed live June 2026.

Openness

2 medium confidence
2.0
platform
closed(hosted SaaS
source
partial(patronus-py is a client wiring to the hosted API, not the engine)
client-SDKs
open(patronus-py public on GitHub, v0.1.25, but only 7 stars and wires to the hosted API)

The open-source SDK is a thin client to a closed hosted evaluation backend; the core scoring engine and the managed Lynx/Glider evaluators are proprietary and not self-hostable. License text is not displayed on the SDK repo page. Class corrected from open_core on 2026-07-30. In this ladder open_core means an OSI core with functionality withheld for a paid tier, and there is no open core here, only an open periphery. A repo that exists while the runtime does not ship is source:partial, which the ladder scores 2/source_available. The score itself was already 2.

Adoption

not assessed

No disclosed user/customer counts or usage volume found on the docs/homepage this run; open source SDK has only 7 GitHub stars (stars_fallback would cap at hobbyist and is not an honest signal for a hosted platform). Declining to assign a level rather than infer from marketing.

Capability

3 medium confidence
3.0

Solid, broad commercial eval feature set (evaluators + experiments + monitoring + agent eval) but not frontier-defining within the category and the closed backend limits reproducibility. Mid-tier on the feature matrix.

Unchanged since 2026-07-30 (last edited, not re-checked)