Evaluating Amazon Connect Customer AI Agents with DeepEval for MRM
Contact Center Blog
This article demonstrates how to build an automated evaluation pipeline for Amazon Connect AI Agents using DeepEval and Amazon Bedrock to meet Model Risk Management (MRM) requirements for regulated financial institutions.
- Deploy a two-layer evaluation pipeline testing both deterministic tool execution and end-to-end conversational behavior
- Configure GEval and AnswerRelevancy metrics backed by Amazon Bedrock running entirely within your AWS account
- Run scored evaluations across nominal, adversarial, ambiguous, and out-of-scope test categories
- Generate structured CSV, JSON, and Markdown reports with quantitative scores and reasoning for MRM documentation
- Integrate into CI/CD pipelines so every agent change automatically produces before-and-after evaluation evidence
- Use LLM-as-judge evaluation to assess semantic correctness, relevancy, and guardrail compliance without manual testing
The pipeline produces auditable, quantitative artifacts required by MRM teams and regulatory frameworks like the Federal Reserve's SR 11-7 guidance, enabling systematic validation of AI agents at scale.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.