Home icon

How Rufus doubled their inference speed and handled Prime Day traffic with AWS AI chips and parallel decoding

Machine Learning Blog



This article details how Rufus, Amazon's AI shopping assistant, optimized its inference performance for Prime Day 2024 using AWS AI chips and parallel decoding techniques.

  • Faced challenge of handling millions of queries per minute with a 300 ms latency requirement
  • Implemented parallel decoding with multiple decoding heads to predict tokens simultaneously
  • Used AWS Trainium and Inferentia chips to accelerate performance
  • Achieved key performance improvements:
    • Two times faster token generation
    • 50% reduction in inference costs
    • Seamless scalability during peak traffic
  • Utilized Neuronx-Distributed Inference (NxDI) framework for implementation

The solution demonstrates how innovative AI chip and decoding techniques can dramatically improve large language model inference performance and efficiency.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Aug 13
2025
How Amazon scaled Rufus by building multi-node inference using AWS Trainium chips and vLLM
Oct 10
2024
Scaling Rufus, the Amazon generative AI-powered conversational shopping assistant with over 80,000 AWS Inferentia and AWS Trainium chips, for Prime Day
Nov 20
2025
How Rufus scales conversational shopping experiences to millions of Amazon customers with Amazon Bedrock
Aug 13
2024
How AWS powered Prime Day 2024 for record-breaking sales

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.