Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts
News
This article announces model caching support in Amazon SageMaker HyperPod, an inference optimization that pre-loads model weights and container images to dramatically reduce cold start times.
- Weights cache stores model weights on local NVMe for fast pod startup instead of network pulls from S3 or FSx
- Image cache pre-pulls container images to eliminate ECR download delays, cutting over two minutes (97% reduction)
- Benchmarks show approximately 60% faster scale-out across models ranging from 57 GB to 145 GB
- Automatic fallback to original source if pod lands on node without warm cache ensures reliability
- Enabled via HyperPod Inference Operator using modelCacheConfig in InferenceEndpointConfig or JumpStartModel
- Generally available in all regions where SageMaker HyperPod is supported
Model caching eliminates inference bottlenecks for LLM workloads like chat assistants, RAG, and document analysis by enabling faster deployments and scale-out events.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 10
2026
2026
Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
May 20
2026
2026
Amazon SageMaker HyperPod now supports data capture for inference workloads
Jul 10
2025
2025
Amazon SageMaker HyperPod launches model deployments to accelerate the generative AI model development lifecycle
Jun 16
2026
2026
Introducing container caching in Amazon SageMaker AI for faster model scaling
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.