AI Potluck
Model components / Evaluation code

Amazon Bedrock Evaluations

Amazon Web Services

Managed evaluation service inside Amazon Bedrock for scoring foundation models and RAG systems. It runs programmatic metrics including BERTScore and F1, a judge model over custom prompt datasets, and human evaluation using either the customer's own workforce or one AWS manages. RAG evaluation covers retrieval and retrieve-and-generate, and bring-your-own inference responses lets external models be scored.

Model evaluation reached general availability in April 2024 and the judge and RAG variants in March 2025. AWS publishes no usage or job-count figure. Verified 2026-08-13 via the Bedrock Evaluations product page and its general-availability announcement.

Openness

1 high confidence
1.0
license
proprietary(AWS managed service)
source
closed
deployment
cloud-only(AWS-hosted eval jobs, S3 I/O)
billing
pay-as-you-go(judge-model inference + AWS-managed human workforce)

Proprietary AWS managed service; no source, no self-host, runs as billed Bedrock evaluation jobs.

Adoption

3 low confidence
3.0

Generally available since 2025-03-20 and distributed inside Amazon Bedrock, a large enterprise AI surface, but AWS publishes no standalone usage count for the Evaluations feature itself. Level 3 rests on the reported traction of the Bedrock platform around it rather than on a measured figure, and confidence is low precisely because the signal belongs to the parent surface. Being a closed service, it offers no star or download count to fall back on.

Capability

4 medium confidence
4.0

A broad managed eval suite covering models, RAG and agents, with judge, programmatic and human grading modes, four families of metrics, and portability through bring-your-own responses. Comprehensive, but a custom-dataset eval product rather than a standardized public benchmark leaderboard, which is why it sits at 4.

Verified 2026-08-13