Home icon

Evaluate healthcare generative AI applications using LLM-as-a-judge on AWS

Machine Learning Blog



The article discusses a novel approach to evaluating healthcare generative AI applications using LLM-as-a-judge on AWS, focusing on generating radiology report impressions using Amazon Bedrock Knowledge Bases and RAG (Retrieval Augmented Generation) techniques.

  • Introduced a comprehensive evaluation framework using five key metrics: correctness, completeness, helpfulness, logical coherence, and faithfulness
  • Used MIMIC Chest X-ray dataset with 91,544 radiology reports for testing
  • Leveraged Amazon Bedrock to compare different generative models like Anthropic's Claude and Amazon Nova
  • Demonstrated high-performance scores across dev1 and dev2 datasets, with correctness and logical coherence metrics reaching 0.98-0.99
  • Provides a systematic method to assess medical AI applications' accuracy, reliability, and clinical utility

The solution represents a significant advancement in maintaining reliability and accuracy of AI-generated medical content, with potential for broader healthcare applications.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Feb 5
2025
Use generative AI on AWS for efficient clinical document analysis
Jun 16
2026
Building a HIPAA-ready generative AI architecture for healthcare on AWS
Jul 17
2025
Evaluating generative AI models with Amazon Nova LLM-as-a-Judge on Amazon SageMaker AI
May 8
2024
How healthcare organizations use generative AI on AWS to turn data into better patient outcomes

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.