Home icon

Evaluating generative AI models with Amazon Nova LLM-as-a-Judge on Amazon SageMaker AI

Machine Learning Blog



This article introduces Amazon Nova LLM-as-a-Judge, a new capability on Amazon SageMaker AI for evaluating generative AI models with advanced reasoning capabilities. The key highlights include:

  • A novel approach to model evaluation using LLMs to compare model outputs
  • Trained on diverse datasets spanning 90+ languages with rigorous bias reduction
  • Provides pairwise comparisons between model iterations with statistical confidence metrics
  • Achieves 45% accuracy on JudgeBench and 68% on PPE benchmarks
  • Integrated with Amazon SageMaker AI for scalable, systematic model assessment

The solution enables organizations to systematically evaluate generative AI model performance beyond traditional metrics, offering a more nuanced and flexible assessment approach that closely reflects human preferences.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Feb 6
2026
Evaluate generative AI models with an Amazon Nova rubric-based LLM judge on Amazon SageMaker AI (Part 2)
Jun 24
2025
Power Your LLM Training and Evaluation with the New SageMaker AI Generative AI Tools
Nov 26
2025
Evaluate models with the Amazon Nova evaluation container using Amazon SageMaker AI
Apr 21
2025
Build an automated generative AI solution evaluation pipeline with Amazon Nova

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.