Home icon

Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1

Machine Learning Blog



This article demonstrates how to deploy a text-to-speech (TTS) model on Amazon SageMaker AI using the AWS vLLM-Omni Deep Learning Container to enable real-time voice applications with streaming audio output.

  • Deploy Qwen3-TTS model using AWS vLLM-Omni DLC for multimodal inference on SageMaker AI
  • Stream text input and audio output over persistent bidirectional HTTP/2 WebSocket connections
  • Use SageMaker instance pools for flexible GPU instance selection (ml.g6.xlarge, ml.g6e.xlarge, ml.g5.xlarge, ml.g4dn.xlarge)
  • Start playing speech before full response generation completes for low-latency voice interactions
  • Includes step-by-step deployment guide, Gradio client application, and reproducible code samples

The vLLM-Omni DLC extends vLLM to support multimodal models for text, audio, images, and video, enabling specialized inference patterns beyond traditional text generation.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

May 20
2026
Build real-time voice applications with Amazon SageMaker AI and vLLM
Sep 28
2026
Generate images and video with vLLM-Omni on SageMaker AI – Part 2
Sep 25
2026
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
Sep 24
2026
Speaker-labeled transcription with WhisperX on SageMaker AI

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.