Home icon

Evaluating Amazon Connect Customer AI Agents with DeepEval for MRM

Contact Center Blog



This article demonstrates how to build an automated evaluation pipeline for Amazon Connect AI Agents using DeepEval and Amazon Bedrock to meet Model Risk Management (MRM) requirements for regulated financial institutions.

  • Deploy a two-layer evaluation pipeline testing both deterministic tool execution and end-to-end conversational behavior
  • Configure GEval and AnswerRelevancy metrics backed by Amazon Bedrock running entirely within your AWS account
  • Run scored evaluations across nominal, adversarial, ambiguous, and out-of-scope test categories
  • Generate structured CSV, JSON, and Markdown reports with quantitative scores and reasoning for MRM documentation
  • Integrate into CI/CD pipelines so every agent change automatically produces before-and-after evaluation evidence
  • Use LLM-as-judge evaluation to assess semantic correctness, relevancy, and guardrail compliance without manual testing

The pipeline produces auditable, quantitative artifacts required by MRM teams and regulatory frameworks like the Federal Reserve's SR 11-7 guidance, enabling systematic validation of AI agents at scale.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Aug 6
2026
Automate customer complaint classification with AI agents on AWS
Aug 24
2026
Amazon Connect Customer now supports information extraction for agent voice and chat conversations
Jul 30
2026
Amazon Connect Customer now automatically finds example agent evaluations for tailored coaching
Jul 21
2026
Amazon Connect Customer launches metrics for agent queues on analytics dashboards

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.