Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
Machine Learning Blog
This article describes an architecture for scaling Mixture-of-Experts (MoE) reinforcement learning workloads on AWS using Amazon EKS, Elastic Fabric Adapter (EFA), and DeepEP, achieving 40% improved throughput.
- MoE models require balancing elastic rollout generation with tightly coupled policy training across heterogeneous compute resources
- Expert Parallelism introduces dynamic all-to-all token routing that becomes communication-bound as jobs scale beyond single instances
- DeepEP over EFA replaces generic all-to-all collectives with specialized dispatch and combining kernels optimized for sparse MoE traffic patterns
- EKS cluster topology separates GPU nodes for training/rollout, CPU nodes for environments, and memory-optimized instances for experience buffers
- Spot Instances reduce rollout generation costs since distributed inference workers tolerate interruptions better than tightly coupled training
- Reference implementation uses CUDA 13.0, PyTorch 2.12.1, NCCL 2.31, EFA 1.49, and DeepEP 2.0 with detailed Dockerfile and deployment examples
The architecture enables efficient large-scale RL training by independently scaling components, balancing resource constraints across compute/memory/bandwidth, and optimizing inter-node communication for MoE workloads.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2025
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.