How Amazon scaled Rufus by building multi-node inference using AWS Trainium chips and vLLM
Machine Learning Blog
This article details how Amazon scaled Rufus, its generative AI shopping assistant, using multi-node inference with AWS Trainium chips and vLLM. Key highlights include:
- Developed a leader/follower multi-node inference architecture to handle large language models
- Used hybrid parallelism strategies to maximize compute and memory utilization across nodes
- Implemented a multi-node inference unit abstraction on Amazon ECS for reliable scaling
- Utilized Elastic Fabric Adapter (EFA) for low-latency, high-bandwidth cross-node communication
- Deployed the solution across tens of thousands of Trainium chips to support Prime Day traffic
The solution enables Amazon to run larger, more capable AI models with improved performance, throughput, and customer engagement for the Rufus shopping assistant.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
May 28
2025
2025
How Rufus doubled their inference speed and handled Prime Day traffic with AWS AI chips and parallel decoding
Oct 10
2024
2024
Scaling Rufus, the Amazon generative AI-powered conversational shopping assistant with over 80,000 AWS Inferentia and AWS Trainium chips, for Prime Day
Mar 30
2026
2026
Accelerate CPU-based AI inference workloads using Intel AMX on Amazon EC2
Apr 15
2026
2026
Accelerating decode-heavy LLM inference with speculative decoding on AWS Trainium and vLLM
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.