AI Potluck
Model components / Evaluation code

Vals AI

Vals AI

Domain-specialist benchmarking service that runs proprietary, held-out evals over finance, law, software, and healthcare tasks, including a private Canadian-court-case QA benchmark and the Vals Index. Picked by enterprises that need closed/private benchmarks because public benchmarks suffer from contamination. Publishes a leaderboard of all major frontier models scored on these private sets.

Vals AI independent LLM evaluation platform; Vals Index + Vals Multimodal Index + domain benchmarks (finance, law, healthcare, coding, math, tax). Verified live June 2026 (leaderboard updated 2026-06-04, includes Claude Opus 4.8, Qwen 3.7 Plus, MiniMax M3). SaaS-only, no public source repo.

Openness

1 high confidence
1.0
license
proprietary
source
closed(SaaS-only,no public repo)
benchmarks
proprietary/private-held-out(Vals Index on non-public datasets)
eval-harness
closed

Closed third-party eval platform; benchmarks deliberately held private to prevent dataset leakage, eval infra not open sourced.

  • https://www.vals.ai recorded 2026-06-04

    proprietary SaaS eval platform, 'proprietary benchmarks' require contact, no GitHub/self-host option

  • https://www.vals.ai/about recorded 2026-06-04

    runs evaluations in-house on private/held-out datasets as neutral third party

Adoption

3 medium confidence
3.0

Widely cited independent leaderboard authority: 24 models tracked on the Vals Index (updated 2026-06-04); benchmark figures (e.g. GLM-5.1) are referenced in the project's own flagship scoring sources and across third-party model coverage; partnered with Legaltech Hub for industry studies. No standalone user/traffic count published, so reported_traction; level 3 reflects broad referenced use as a benchmark authority rather than a measured usage figure.

  • https://www.vals.ai/benchmarks recorded 2026-06-04

    live multi-domain leaderboards (Vals Index, Multimodal Index) tracking 24+ frontier models, updated June 2026

Capability

4 medium confidence
4.0

Strong feature_matrix: leading independent/blind benchmark suite with deep vertical coverage and agentic-eval, but narrower than a general harness (no self-host, fixed proprietary task set). Not a frontier-definer (5) on standardization since methodology is closed; placed at 4.

  • https://www.vals.ai recorded 2026-06-04

    Vals Index + Multimodal Index + domain benchmarks (finance/law/healthcare/coding/math/education) incl TaxEval, CorpFin, MedCode, LegalBench

Unchanged since 2026-06-24 (last edited, not re-checked)