AI Potluck
Model components / Evaluation code

Vertex AI Gen AI Evaluation Service

Google Cloud

Managed evaluation service on Vertex AI (now part of Gemini Enterprise Agent Platform) that scores generative-AI outputs using three approaches: computation-based metrics (ROUGE, BLEU, exact match), LLM-as-judge with Gemini as the judge model (including adaptive rubrics that auto-generate per-prompt pass/fail tests), and AutoSxS pairwise pipelines that A/B test two models with a judge declaring a winner. Differentiates on tight integration with Vertex AI pipelines, Model Garden, Gemini tuning workflows, and LiteLLM for third-party models; eval datasets can be built from production logs or synthetic data and run across seven global regions. Used by Vertex AI customers including Wendy's, Wayfair, Verizon, Mercedes-Benz, and Etsy (per Google Cloud case studies) as the canonical eval surface for Gemini-based applications.

Google Cloud Vertex AI managed Gen AI evaluation service; evaluates models, agents, and RAG. Computation-based metrics, model-based/autorater (managed rubric-based + custom judge models), pointwise + pairwise (AutoSxS). Part of Gemini Enterprise Agent Platform. Verified live June 2026 (docs current).

Openness

1 high confidence
1.0
license
proprietary(GCP managed service)
source
closed
deployment
cloud-only(Vertex AI pipelines/AutoSxS)
billing
GCP managed-service pricing

Proprietary Google Cloud managed service; no source, no self-host, runs as Vertex AI eval pipelines.

Adoption

not assessed

Managed GCP feature with no published standalone usage/customer count and no download or stars signal (closed cloud service). Unlike Bedrock Evaluations there was no clear GA-scale traction figure verifiable this run, so declining to assign a level rather than infer from the parent Vertex surface.

Capability

4 medium confidence
4.0

Broad managed eval suite: model/agent/RAG targets, both computation- and autorater-based metrics, and pointwise + pairwise (AutoSxS) comparison. Comparable scope to Bedrock Evaluations; a custom-dataset eval product rather than a public standardized leaderboard, so 4 not 5.

Unchanged since 2026-06-09 (last edited, not re-checked)