Amazon Bedrock Evaluations
Amazon Web ServicesManaged evaluation service inside Amazon Bedrock for scoring foundation models and RAG systems. It runs programmatic metrics including BERTScore and F1, a judge model over custom prompt datasets, and human evaluation using either the customer's own workforce or one AWS manages. RAG evaluation covers retrieval and retrieve-and-generate, and bring-your-own inference responses lets external models be scored.
Model evaluation reached general availability in April 2024 and the judge and RAG variants in March 2025. AWS publishes no usage or job-count figure. Verified 2026-08-13 via the Bedrock Evaluations product page and its general-availability announcement.
Openness
1 high confidence- license
- proprietary(AWS managed service)
- source
- closed
- deployment
- cloud-only(AWS-hosted eval jobs, S3 I/O)
- billing
- pay-as-you-go(judge-model inference + AWS-managed human workforce)
Proprietary AWS managed service; no source, no self-host, runs as billed Bedrock evaluation jobs.
- https://docs.aws.amazon.com/bedrock/latest/userguide/evaluation.html recorded 2026-06-04
managed Bedrock service running automatic/judge/human eval jobs on AWS, CORS S3 I/O, CloudTrail events
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-06-04
AWS-hosted feature, AWS-managed human evaluator workforce option, directs to pricing page
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-08-13
An AWS-hosted evaluation capability inside Amazon Bedrock, offering LLM-as-a-Judge, programmatic and human evaluation, with AWS able to 'engage a team managed by AWS to perform the human evaluation'. The page links to Bedrock pricing. No source code, repository or self-host option is mentioned anywhere on it, and no license is offered.
Adoption
3 low confidenceGenerally available since 2025-03-20 and distributed inside Amazon Bedrock, a large enterprise AI surface, but AWS publishes no standalone usage count for the Evaluations feature itself. Level 3 rests on the reported traction of the Bedrock platform around it rather than on a measured figure, and confidence is low precisely because the signal belongs to the parent surface. Being a closed service, it offers no star or download count to fall back on.
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-06-04
GA announcement (2025-03-20) of model + RAG evaluations as a broadly available Bedrock feature
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-08-13
The GA announcement for model and RAG evaluations carries no usage, customer or job-count figure.
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-08-13
No published usage or customer count for the Evaluations feature.
Capability
4 medium confidenceA broad managed eval suite covering models, RAG and agents, with judge, programmatic and human grading modes, four families of metrics, and portability through bring-your-own responses. Comprehensive, but a custom-dataset eval product rather than a standardized public benchmark leaderboard, which is why it sits at 4.
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-06-04
model eval (LLM-as-judge/programmatic/human) + RAG eval (retrieval, retrieve-and-generate) with safety/quality metrics
- https://aws.amazon.com/bedrock/evaluations/ recorded 2026-08-13
LLM-as-a-Judge over custom prompt datasets, programmatic metrics including BERTScore and F1, human evaluation with your own or an AWS-managed workforce, and both retrieval and retrieve-and-generate RAG evaluation.
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-08-13
The four metric areas (quality, user experience, instruction compliance, safety) and the citation precision and coverage metrics are documented.
- https://aws.amazon.com/blogs/machine-learning/evaluate-models-or-rag-systems-using-amazon-bedrock-evaluations-now-generally-available/ recorded 2026-06-04
four metric areas (quality/UX/instruction-compliance/safety) and citation precision/coverage metrics
Verified 2026-08-13