Speaker-labeled transcription with WhisperX on SageMaker AI
Machine Learning Blog
This article demonstrates how to deploy the AWS WhisperX Deep Learning Container on Amazon SageMaker AI for production speech-to-text workloads with speaker diarization and precise timestamps.
- WhisperX extends OpenAI's Whisper with per-word timestamps, speaker labels, and faster batched inference for accurate transcription
- Deploy as real-time endpoint for short interactive clips (under 60 seconds) or asynchronous endpoint for long audio without time limits
- Real-time endpoint returns transcripts inline; asynchronous endpoint brokers input/output through S3 and supports polling
- Must pin GPU AMI version to al2-ami-sagemaker-inference-gpu-3-1 to avoid container startup failures
- Scale by adding instances with MaxConcurrentInvocationsPerInstance=1, not by increasing concurrency per container
- Supports multiple output formats (JSON, verbose JSON, SRT, VTT) for analytics pipelines and video editors
- Asynchronous endpoints can autoscale to zero for cost savings on bursty batch workloads
WhisperX on SageMaker AI enables accurate, searchable transcripts with speaker identification for contact centers, media, legal, and healthcare applications.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 25
2026
2026
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
Sep 28
2026
2026
Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
Sep 15
2026
2026
Build an AI-powered product tagging system with Amazon SageMaker serverless model customization
Sep 16
2024
2024
Whisper audio transcription powered by AWS Batch and AWS Inferentia
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.