Amazon SageMaker AI Now Supports Capacity-Aware Inference with Automatic Instance Fallback
News
This article announces capacity-aware inference with automatic instance fallback for Amazon SageMaker AI endpoints.
- SageMaker AI automatically provisions from prioritized instance types when preferred capacity unavailable
- Supports Single Model Endpoints, InferenceComponent-based endpoints, and Asynchronous Inference endpoints
- Scales down by removing lowest-priority instances first, preserving preferred infrastructure
- Specify different optimized models per instance type or use SageMaker inference recommendations
- Per-instance-type CloudWatch metrics provide visibility into latency, throughput, and GPU utilization
- Available in 16 AWS regions globally
SageMaker AI now handles capacity constraints automatically, ensuring reliable endpoint creation and scaling without manual intervention.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
May 4
2026
2026
Capacity-aware inference: Automatic instance fallback for SageMaker AI endpoints
Apr 22
2026
2026
Amazon SageMaker AI now supports optimized generative AI inference recommendations
Apr 22
2026
2026
Amazon SageMaker AI launches optimized generative AI inference recommendations
May 21
2026
2026
Amazon SageMaker AI now supports OpenAI-compatible APIs for inference endpoints
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.