Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
Machine Learning Blog
This article demonstrates how to deploy the Qwen3-TTS-12Hz-1.7B-Base text-to-speech model from Amazon SageMaker JumpStart to create real-time voice cloning capabilities.
- Voice cloning generates speech in a target speaker's voice from short reference recordings without model retraining
- Supports 10 languages with cross-lingual cloning to preserve speaker identity across languages
- Deploy using SageMaker JumpStart with ml.g6.4xlarge GPU instance (24 GB L4) and GPU memory utilization set to 0.45
- Model uses two-stage architecture: talker stage generates speech tokens, code2wav stage renders waveforms
- Invoke endpoint with base64-encoded reference audio, transcript, target text, and custom routing attribute
- Monitor with CloudWatch metrics including GPU utilization, model latency, and concurrent requests
- Supports use cases: content localization, customer experience, e-learning, creative prototyping, and conversational AI
Self-hosted voice cloning on SageMaker provides cost efficiency, data control, and scalability without managing underlying infrastructure.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 28
2026
2026
Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
Sep 24
2026
2026
Speaker-labeled transcription with WhisperX on SageMaker AI
May 20
2026
2026
Build real-time voice applications with Amazon SageMaker AI and vLLM
Sep 15
2026
2026
Build an AI-powered product tagging system with Amazon SageMaker serverless model customization
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.