Amazon Bedrock Evaluations
Amazon Web ServicesManaged evaluation service inside Amazon Bedrock for scoring foundation models and RAG systems with three modes: automatic (BERTScore, F1, accuracy, robustness, toxicity), LLM-as-a-Judge (correctness, completeness, harmfulness, custom prompts), and human evaluation (internal workforce or AWS-managed reviewer team). Differentiates by being native to the Bedrock model catalog and Knowledge Bases; eval jobs run against any Bedrock-supported model, results tie directly into the model deployment pipeline, and comparative analysis across runs is built in. Bedrock model evaluation went GA April 2024; the LLM-as-a-Judge variant went GA March 2025. Bedrock as a platform serves tens of thousands of customers (per AWS, Dec 2024) including adidas, Intuit, Pfizer, Salesforce, Siemens, Thomson Reuters, and the NYSE.
AWS managed evaluation feature within Amazon Bedrock; model eval (LLM-as-a-judge, programmatic incl BERTScore/F1, human), RAG eval (retrieval + retrieve-and-generate). RAG eval + LLM-as-a-judge GA 2025-03-20 (preview at re:Invent Dec 2024); 'bring your own inference responses' supports external/on-prem models. Verified live June 2026.
Openness
1 high confidence- license
- proprietary(AWS managed service)
- source
- closed
- deployment
- cloud-only(AWS-hosted eval jobs, S3 I/O)
- billing
- pay-as-you-go(judge-model inference + AWS-managed human workforce)
Proprietary AWS managed service; no source, no self-host, runs as billed Bedrock evaluation jobs.
- https://docs.aws.amazon.com/bedrock/latest/userguide/evaluation.html recorded 2026-06-04
managed Bedrock service running automatic/judge/human eval jobs on AWS, CORS S3 I/O, CloudTrail events
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-06-04
AWS-hosted feature, AWS-managed human evaluator workforce option, directs to pricing page
Adoption
3 low confidenceGA since 2025-03-20 and distributed inside Amazon Bedrock (a large enterprise AI surface), but AWS publishes no standalone usage count for the Evaluations feature specifically. Level 3 = reported traction via the Bedrock platform footprint, not a measured figure; low confidence because the signal is the parent surface, not the feature. No stars fallback (closed service).
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-06-04
GA announcement (2025-03-20) of model + RAG evaluations as a broadly available Bedrock feature
Capability
4 medium confidenceBroad managed eval suite covering model + RAG + agent, judge/programmatic/human modes, and four metric families with BYO-responses portability. Comprehensive but a custom-dataset eval product, not a standardized public benchmark leaderboard; placed at 4.
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-06-04
model eval (LLM-as-judge/programmatic/human) + RAG eval (retrieval, retrieve-and-generate) with safety/quality metrics
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-06-04
four metric areas (quality/UX/instruction-compliance/safety) and citation precision/coverage metrics
Unchanged since 2026-06-09 (last edited, not re-checked)