Use custom metrics to evaluate your generative AI application with Amazon Bedrock
Machine Learning Blog
This article introduces custom metrics for evaluating generative AI applications using Amazon Bedrock Evaluations. The key features and highlights include:
- Ability to create custom evaluation metrics for model and RAG system assessments
- Support for both numerical and categorical scoring systems
- Flexible metric creation with template variables like {{prompt}} and {{prediction}}
- Built-in templates and options to create metrics from scratch
- Capability to save and reuse custom metrics across evaluation jobs
The new feature allows organizations to define evaluation criteria specific to their business requirements, extending the LLM-as-a-judge framework and enabling more meaningful AI system performance assessments.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2025
2025
2025
2024
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.