Agentic conversational video intelligence built on AWS
Machine Learning Blog
This article demonstrates how to build a conversational video intelligence system using agentic AI on AWS that answers natural language questions about video content by orchestrating Bedrock, Rekognition, and Transcribe services.
- Agentic architecture uses an LLM to decide at runtime which AWS services to invoke based on user queries
- Supports transcription-based queries, visual analysis, face matching, and comprehensive video summaries
- Caches analysis results for fast follow-up queries (under 1 second) while initial analysis takes 5-10 minutes
- Extends beyond video to documents, clinical notes, and other modalities by adding new tool functions
- Deploys on ECS Fargate with CloudFront, Cognito authentication, per-user S3 isolation, and threat modeling
- Reduces manual video review time by approximately 80% compared to traditional manual analysis
The agentic pattern eliminates fixed processing pipelines, enabling flexible multi-modal AI assistants that scale to new capabilities by adding tool functions without workflow logic changes.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.