AI Potluck
Model components / Evaluation code

Amazon Bedrock Evaluations

Amazon Web Services

Managed evaluation service inside Amazon Bedrock for scoring foundation models and RAG systems with three modes: automatic (BERTScore, F1, accuracy, robustness, toxicity), LLM-as-a-Judge (correctness, completeness, harmfulness, custom prompts), and human evaluation (internal workforce or AWS-managed reviewer team). Differentiates by being native to the Bedrock model catalog and Knowledge Bases; eval jobs run against any Bedrock-supported model, results tie directly into the model deployment pipeline, and comparative analysis across runs is built in. Bedrock model evaluation went GA April 2024; the LLM-as-a-Judge variant went GA March 2025. Bedrock as a platform serves tens of thousands of customers (per AWS, Dec 2024) including adidas, Intuit, Pfizer, Salesforce, Siemens, Thomson Reuters, and the NYSE.

AWS managed evaluation feature within Amazon Bedrock; model eval (LLM-as-a-judge, programmatic incl BERTScore/F1, human), RAG eval (retrieval + retrieve-and-generate). RAG eval + LLM-as-a-judge GA 2025-03-20 (preview at re:Invent Dec 2024); 'bring your own inference responses' supports external/on-prem models. Verified live June 2026.

Openness

1 high confidence
1.0
license
proprietary(AWS managed service)
source
closed
deployment
cloud-only(AWS-hosted eval jobs, S3 I/O)
billing
pay-as-you-go(judge-model inference + AWS-managed human workforce)

Proprietary AWS managed service; no source, no self-host, runs as billed Bedrock evaluation jobs.

Adoption

3 low confidence
3.0

GA since 2025-03-20 and distributed inside Amazon Bedrock (a large enterprise AI surface), but AWS publishes no standalone usage count for the Evaluations feature specifically. Level 3 = reported traction via the Bedrock platform footprint, not a measured figure; low confidence because the signal is the parent surface, not the feature. No stars fallback (closed service).

Capability

4 medium confidence
4.0

Broad managed eval suite covering model + RAG + agent, judge/programmatic/human modes, and four metric families with BYO-responses portability. Comprehensive but a custom-dataset eval product, not a standardized public benchmark leaderboard; placed at 4.

Unchanged since 2026-06-09 (last edited, not re-checked)