Home icon

Fault tolerant distributed training on Amazon EKS using NVRx

Machine Learning Blog



This article demonstrates how to integrate NVIDIA Resiliency Extension (NVRx) with PyTorch FSDP training on Amazon EKS to achieve fault-tolerant distributed training with async checkpointing and fast recovery.

  • Async checkpointing overlaps I/O with training, achieving 99%+ efficiency versus 57-61% for synchronous checkpointing
  • In-process restart recovers from soft faults (exceptions, NCCL hangs) in ~10 seconds without container restarts
  • ft_launcher handles hard faults (SIGKILL, OOM) by respawning workers and reloading from latest checkpoint
  • Benchmarks on H100 GPUs show NVRx in-process restart achieves 31% training goodput versus 11.5% for baseline Kubernetes recovery
  • Solution uses Amazon EKS with p5.48xlarge nodes, EFA networking, and FSx for Lustre shared storage

NVRx enables aggressive checkpointing and rapid fault recovery, maximizing GPU utilization for large-scale distributed training workloads.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Sep 25
2026
Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
Aug 12
2026
Hybrid ML inferencing on Amazon EKS with Amazon FSx for NetApp ONTAP and on-premises NetApp
Sep 1
2026
Fast model loading for AI inference on Amazon EKS
Oct 15
2025
Configure and verify a distributed training cluster with AWS Deep Learning Containers on Amazon EKS

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.