Hybrid ML inferencing on Amazon EKS with Amazon FSx for NetApp ONTAP and on-premises NetApp
Storage Blog
This article demonstrates how to run stateful ML inference on Amazon EKS using FSx for NetApp ONTAP for persistent model caching, eliminating cold starts and enabling hybrid cloud deployments.
- FSx for ONTAP provides persistent shared storage for ML model weights, tokenizer files, and compiled GPU kernels across pod restarts and scaling events
- NetApp SnapMirror enables block-level replication of trained models from on-premises ONTAP systems to FSx for ONTAP without manual transfers
- Trident CSI driver automates Kubernetes-native provisioning, snapshots, and data protection for FSx for ONTAP volumes
- Warm start inference pods skip model download and CUDA compilation, reducing startup time from 228 seconds (cold) to ~101 seconds (warm)
- Multi-AZ FSx for ONTAP deployments provide automatic failover and disaster recovery for model artifacts
- Solution includes Terraform automation for EKS cluster, FSx for ONTAP, Karpenter GPU node pools, and Trident deployment
FSx for ONTAP eliminates inference cold starts through persistent caching while enabling clean hybrid architectures where models train on-premises and serve in AWS.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 1
2026
2026
Fast model loading for AI inference on Amazon EKS
Sep 3
2026
2026
AWS Transform announces general availability of Amazon FSx for NetApp ONTAP support
Aug 5
2026
2026
Oracle Machine Learning for SQL on Amazon RDS: Build machine learning models entirely in SQL
Aug 3
2026
2026
Introducing Apache Spark troubleshooting agent for Amazon EMR on EKS
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.