Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
Machine Learning Blog
This article demonstrates how to deploy a text-to-speech (TTS) model on Amazon SageMaker AI using the AWS vLLM-Omni Deep Learning Container to enable real-time voice applications with streaming audio output.
- Deploy Qwen3-TTS model using AWS vLLM-Omni DLC for multimodal inference on SageMaker AI
- Stream text input and audio output over persistent bidirectional HTTP/2 WebSocket connections
- Use SageMaker instance pools for flexible GPU instance selection (ml.g6.xlarge, ml.g6e.xlarge, ml.g5.xlarge, ml.g4dn.xlarge)
- Start playing speech before full response generation completes for low-latency voice interactions
- Includes step-by-step deployment guide, Gradio client application, and reproducible code samples
The vLLM-Omni DLC extends vLLM to support multimodal models for text, audio, images, and video, enabling specialized inference patterns beyond traditional text generation.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.