Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Machine Learning Blog
This article describes how to fine-tune LLM-powered search agents using Amazon SageMaker AI multi-turn reinforcement learning (MTRL) to optimize multi-turn interactions and improve retrieval quality.
- MTRL trains agents across full multi-turn trajectories rather than single steps, optimizing interdependent decisions in search workflows
- Modular agent-environment interface enables custom rewards, tool loops, and conversation shapes with low-code integration
- Serverless execution with per-token pricing eliminates GPU cluster provisioning and management overhead
- Fine-tuned Qwen3.6-27B model achieved +23.7% nDCG@10 gain on BrowseComp-Plus and reduced failure rate from 22.89% to 0.68%
- Minimal configuration required: only three hyperparameters changed from defaults (max_epochs, global_batch_size, rollout_max_concurrency)
- Trajectory and reward observability via MLflow enables inspection of agent behavior turn-by-turn across training steps
Amazon SageMaker AI MTRL provides a practical, low-barrier approach to training specialized search agents with improved reliability and efficiency without deep RL expertise.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.