How to: Use AudioShake to improve accuracy in localization, transcription, and captioning workflows on AWS
Media Blog
The article discusses how to use AudioShake, a sound separation technology, to improve the accuracy of localization, transcription, and captioning workflows on AWS.
Specifically, the article covers:
- An introduction to AudioShake and its ability to isolate dialogue and music/effects from audio tracks, improving downstream processes like ASR, captioning, and dubbing
- A walkthrough of an AWS serverless architecture to separate dialogue and music/effects from media assets using AudioShake Groovy
- An example workflow demonstrating how the separated dialogue audio can be used with AWS services like Amazon Transcribe, Amazon Translate, and Amazon Polly to generate captions, subtitles, and synthesized speech in multiple languages
- Conclusion highlighting AudioShake's ability to improve speech recognition accuracy through automated audio refinement on AWS
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
May 29
2025
2025
Automating audio editing and transcoding using AWS
May 28
2025
2025
Enhanced Performance for Whisper Audio Transcription on AWS Batch and AWS Inferentia
Mar 5
2025
2025
CaptionHub Live on AWS expands the reach of live video with subtitling
Apr 1
2026
2026
Live streaming localization and accessibility using AWS Media Services
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.