Home icon

Optimizing Mobileye’s REM™ with AWS Graviton: A focus on ML inference and Triton integration

Machine Learning Blog



This article describes how Mobileye optimized its Road Experience Management (REM) system's change detection pipeline using AWS Graviton processors and Triton Inference Server.

  • Chose CPU over GPU instances for cost efficiency despite slower inference performance
  • Implemented Triton Inference Server to centralize model serving, reducing memory per task from 8.5GB to 2.5GB
  • Optimized Triton Docker image from 15GB to 2.7GB by removing unnecessary backends
  • Added AWS Graviton instances to diversify compute pool and improve Spot Instance availability
  • Achieved over 2x throughput improvement through instance diversification and Graviton performance
  • Graviton instances delivered 19.4 samples/second versus 13.5 for comparable non-Graviton CPUs

Mobileye successfully optimized their change detection system by prioritizing cost efficiency, centralizing inference, and leveraging AWS Graviton's price-performance advantages for ML workloads.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Nov 25
2025
Warner Bros. Discovery achieves 60% cost savings and faster ML inference with AWS Graviton
Nov 24
2025
How potential performance upside with AWS Graviton helps reduce your costs further
Nov 26
2025
Enhancing and monitoring network performance when running ML Inference on Amazon EKS
Dec 17
2025
Improve Aurora PostgreSQL throughput by up to 165% and price-performance ratio by up to 120% using Optimized Reads on AWS Graviton4-based R8gd instances

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.