AI Potluck
Model components / Evaluation code

Vertex AI Gen AI Evaluation Service

Google Cloud

Managed evaluation service on Vertex AI, part of the Gemini Enterprise Agent Platform, that scores generative-AI outputs three ways: computation-based metrics such as ROUGE and BLEU, model-based autorater metrics with adaptive rubrics that generate per-prompt pass/fail tests, and AutoSxS pairwise pipelines that A/B test two models with a judge declaring a winner. It covers models, agents and RAG grounding, and evaluation datasets can be built from production logs or synthetic data.

Google publishes no usage, customer or run-count figure for the service. Verified 2026-08-13 via the Vertex AI evaluation overview documentation.

Openness

1 high confidence
1.0
license
proprietary(GCP managed service)
source
closed
deployment
cloud-only(Vertex AI pipelines/AutoSxS)
billing
GCP managed-service pricing

Proprietary Google Cloud managed service; no source, no self-host, runs as Vertex AI eval pipelines. The open-model references on the overview are about models the service can deploy and evaluate, not about the service itself, which is offered with no source, no self-host and no license.

Adoption

not assessed

A managed Google Cloud feature with no published standalone usage or customer count, and no download or star signal to fall back on, being a closed cloud service. Unlike Amazon Bedrock Evaluations it offers no GA-scale traction figure that can be checked, so no level is assigned rather than inferring one from the size of the parent Vertex AI surface.

Capability

4 medium confidence
4.0

A broad managed eval suite: model, agent and RAG targets, both computation-based and autorater metrics, and pointwise as well as pairwise (AutoSxS) comparison. Comparable in scope to Amazon Bedrock Evaluations, which sits at the same level, and like it a custom-dataset eval product rather than a public standardized leaderboard, so 4 rather than 5.

Verified 2026-08-13