How Amazon Achieved Full Stack Observability Across 400 Offices with Amazon OpenSearch Serverless
AWS Cloud Operations Blog
This article describes how Amazon Corporate Infrastructure Services built a Full Stack Observability platform using Amazon OpenSearch Serverless to monitor 330,000+ employees across 400 offices globally.
- Consolidated fragmented monitoring from domain-specific tools into unified platform eliminating data silos
- Achieved 5-minute Mean Time to Detect (MTTD) target, down from 30-60 minutes baseline
- Used Amazon OpenSearch Ingestion for data normalization, transformation, and enrichment across multiple telemetry sources
- Implemented three-layer architecture: data sources, processing/storage, and insights/alerting with OpenSearch Dashboards
- Scaled from 3 pilot sites to 24 sites in Phase 1 using repeatable ingestion patterns and Infrastructure as Code
- Achieved 99.9% platform uptime and 220% projected annual ROI through engineering time and productivity savings
- Key lessons: start with business outcomes, embrace open standards (OpenTelemetry), design for scale, invest in data quality, and enable self-service dashboards
The serverless approach eliminated operational complexity while enabling rapid expansion and proactive issue detection across Amazon's global infrastructure.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2024
2024
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.