Home icon

Generate training data and cost-effectively train categorical models with Amazon Bedrock

Machine Learning Blog



This article explores using Amazon Bedrock and generative AI to generate high-quality labeled training data for multiclass classification models, specifically focusing on categorizing customer support cases.

  • Traditional data labeling is expensive, time-consuming, and prone to class imbalances
  • Prompt engineering with XML tags can guide large language models like Claude 3.5 Sonnet to generate accurate labeled datasets
  • The approach achieved over 90% accuracy in categorizing support case root causes into six categories:
    • Billing Inquiry
    • Security Awareness
    • Feature Request
    • Software Defect
    • Documentation Improvement
    • Customer Education
  • Key steps include:
    • Defining clear category descriptions
    • Providing good and bad example classifications
    • Using XML-structured prompts
    • Iteratively refining the prompt
  • The generated labeled data can be used to train machine learning models using AutoGluon

The method offers a cost-effective and accurate alternative to manual data labeling for complex classification tasks.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Apr 8
2025
Build an enterprise synthetic data strategy using Amazon Bedrock
Apr 22
2025
Optimizing cost for using foundational models with Amazon Bedrock
Mar 25
2025
Evaluate and improve performance of Amazon Bedrock Knowledge Bases
Mar 20
2025
Unleashing the multimodal power of Amazon Bedrock Data Automation to transform unstructured data into actionable insights

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.