Real-time voice agents with Stream Vision Agents and Amazon Nova 2 Sonic
Machine Learning Blog
This article explains how to build production-ready real-time voice agents by combining Stream's Vision Agents framework with Amazon Nova 2 Sonic and Amazon Bedrock.
- Amazon Nova 2 Sonic provides speech-to-speech AI with native turn detection and function calling capabilities
- Stream's Vision Agents abstracts infrastructure complexity with plugin-based architecture and global edge network
- Sub-500ms end-to-end latency achieved through distributed SFU media forwarding and optimized audio streaming
- Function calling enables agents to query databases, call APIs, and trigger workflows during conversations
- Complete working agent built in under 30 lines of Python code using decorators and async patterns
- Supports custom pipelines with separate STT/TTS providers like Deepgram and Cartesia
- Amazon Polly TTS integration available for non-realtime LLM architectures
- Ideal for no-screen environments (driving, field service) and high-volume phone support automation
Vision Agents simplifies building voice AI applications by handling connection management, audio encoding, and orchestration, allowing developers to focus on agent logic and user experience.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.