Optimizing Mobileye’s REM™ with AWS Graviton: A focus on ML inference and Triton integration
Machine Learning Blog
This article describes how Mobileye optimized its Road Experience Management (REM) system's change detection pipeline using AWS Graviton processors and Triton Inference Server.
- Chose CPU over GPU instances for cost efficiency despite slower inference performance
- Implemented Triton Inference Server to centralize model serving, reducing memory per task from 8.5GB to 2.5GB
- Optimized Triton Docker image from 15GB to 2.7GB by removing unnecessary backends
- Added AWS Graviton instances to diversify compute pool and improve Spot Instance availability
- Achieved over 2x throughput improvement through instance diversification and Graviton performance
- Graviton instances delivered 19.4 samples/second versus 13.5 for comparable non-Graviton CPUs
Mobileye successfully optimized their change detection system by prioritizing cost efficiency, centralizing inference, and leveraging AWS Graviton's price-performance advantages for ML workloads.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2025
2025
2025
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.