Vals AI
Vals AIIndependent benchmarking service that runs held-out evaluations over finance, law, healthcare, coding, math, education, cybersecurity and games, and publishes composite indices alongside a leaderboard of frontier models scored on those private sets. The benchmarks stay private specifically because public ones suffer contamination.
The evaluation runtime is hosted and the held-out datasets do not ship, though Vals has released the Valkyrie agentic-benchmark runner and its model library. Verified 2026-08-14 via the Vals AI about page and the two repositories' metadata.
Openness
2 high confidence- license
- proprietary(the Vals Index platform and its held-out sets)+AGPL-3.0(OSI, vals-ai/Valkyrie)+MIT(OSI, vals-ai/model-library)
- source
- partial(harness repos public, Index runtime and datasets do not ship)
- benchmarks
- proprietary/private-held-out(Vals Index on non-public datasets)
- eval-harness
- published(Valkyrie AGPL-3.0 + model-library MIT)
A third-party eval platform whose benchmarks are deliberately held private to prevent dataset leakage, with part of the infrastructure now published. The about page states 'We have built, and even open-sourced, the infrastructure we rely on to run these evaluations reproducibly and at scale across labs. This includes Valkyrie, a distributed system to run agentic benchmarks, and our model library, a free standard API to call models with'. Both repos are live, public and actively pushed - vals-ai/Valkyrie under AGPL-3.0 and vals-ai/model-library under MIT - so neither the source nor the eval harness is entirely closed. What is closed is the part the tier turns on: scores are still 'based on our privately held test sets', the Vals Index runtime is still a hosted service with no self-hostable implementation, and the held-out datasets do not ship. A repo that exists while the runtime does not is only a partial source, which is source-available rather than closed.
- https://www.vals.ai recorded 2026-06-04
proprietary SaaS eval platform, 'proprietary benchmarks' require contact, no GitHub/self-host option
- https://www.vals.ai/about recorded 2026-06-04
runs evaluations in-house on private/held-out datasets as neutral third party
- https://www.vals.ai/about recorded 2026-08-13
Vals describes itself as 'the independent evaluator of artificial intelligence' running evaluations in-house on 'privately held test sets'. It ALSO states it has open-sourced Valkyrie, a distributed system for running agentic benchmarks, and a model library offering a free standard API for calling models, both linked to GitHub.
- https://www.vals.ai/about recorded 2026-08-14
'We have built, and even open-sourced, the infrastructure we rely on to run these evaluations reproducibly and at scale across labs. This includes Valkyrie, a distributed system to run agentic benchmarks, and our model library, a free standard API to call models with.' Beside it, 'Scores are based on our privately held test sets to preserve the integrity and signal of our results' - the harness ships, the sets and the Index runtime do not.
- https://api.github.com/repos/vals-ai/valkyrie recorded 2026-08-14
Repo metadata for vals-ai/Valkyrie - license spdx_id AGPL-3.0, public, not archived, 31 stars, pushed 2026-08-14, described as 'Scalable, cloud-native infrastructure for evaluating AI agents across any benchmark'. A live published harness rather than an abandoned mirror.
- https://api.github.com/repos/vals-ai/model-library recorded 2026-08-14
Repo metadata for vals-ai/model-library - license spdx_id MIT, public, not archived, 21 stars, pushed 2026-07-30, 'Simple provider agnostic LLM gateway'.
Adoption
3 medium confidenceA widely cited independent leaderboard authority. The benchmarks page tracks 43 models across the Vals Index, an RSI Index, a Multimodal Index and legal, finance, healthcare, math, academic, education and coding families; its figures - GLM-5.1, for instance - are referenced in this project's own flagship scoring sources and across third-party model coverage; and it has partnered with Legaltech Hub on industry studies. Vals publishes no user, traffic or customer count of any kind, so level 3 reflects broad referenced use as a benchmark authority rather than a measured usage figure.
- https://www.vals.ai/benchmarks recorded 2026-06-04
live multi-domain leaderboards (Vals Index, Multimodal Index) tracking 24+ frontier models, updated June 2026
- https://www.vals.ai/benchmarks recorded 2026-08-13
43 models tested; Vals Index, RSI Index and Vals Multimodal Index plus legal, finance, healthcare, math, academic, education, coding, cybersecurity and games benchmarks. No usage, traffic or customer figure is published on the page.
Capability
4 medium confidenceA leading independent blind-benchmark suite with deep vertical coverage and agentic evaluation, but narrower than a general harness: there is no self-hosting, and the task set is fixed and proprietary. Not a frontier-definer, because a closed methodology cannot become a standard others reproduce. The vertical coverage holds and has grown, with cybersecurity and games benchmarks added; two Vals repos have since been published, but self-hosting the Index is still not offered, so the assessment is unchanged.
- https://www.vals.ai recorded 2026-06-04
Vals Index + Multimodal Index + domain benchmarks (finance/law/healthcare/coding/math/education) incl TaxEval, CorpFin, MedCode, LegalBench
- https://www.vals.ai recorded 2026-08-13
Vals Index and Vals Multimodal Index over finance (CorpFin v2, Finance Agent v2, TaxEval v2), legal (LegalBench, Case Law v2, Legal Research Bench), healthcare (MedCode, MedScribe, MedQA), coding (SWE-bench Verified, Terminal-Bench, Vibe Code Bench, IOI, LiveCodeBench), math (ProofBench, AIME, GPQA Diamond) and education (SAGE, MMLU Pro, MMMU). 'We run all of our own evaluations and create many of our benchmarks in-house'.
Verified 2026-08-13