Home icon

Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

Machine Learning Blog



This article describes how to fine-tune LLM-powered search agents using Amazon SageMaker AI multi-turn reinforcement learning (MTRL) to optimize multi-turn interactions and improve retrieval quality.

  • MTRL trains agents across full multi-turn trajectories rather than single steps, optimizing interdependent decisions in search workflows
  • Modular agent-environment interface enables custom rewards, tool loops, and conversation shapes with low-code integration
  • Serverless execution with per-token pricing eliminates GPU cluster provisioning and management overhead
  • Fine-tuned Qwen3.6-27B model achieved +23.7% nDCG@10 gain on BrowseComp-Plus and reduced failure rate from 22.89% to 0.68%
  • Minimal configuration required: only three hyperparameters changed from defaults (max_epochs, global_batch_size, rollout_max_concurrency)
  • Trajectory and reward observability via MLflow enables inspection of agent behavior turn-by-turn across training steps

Amazon SageMaker AI MTRL provides a practical, low-barrier approach to training specialized search agents with improved reliability and efficiency without deep RL expertise.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Jun 3
2026
Amazon SageMaker AI launches multi-turn reinforcement learning for AI agent model customization
Jul 2
2026
Best practices for multi-turn reinforcement learning in Amazon SageMaker AI
Sep 23
2026
Query unstructured data in Amazon SageMaker Catalog using generative AI
Jul 11
2025
Advanced fine-tuning methods on Amazon SageMaker AI

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.