Secure short-term GPU capacity for ML workloads with EC2 Capacity Blocks for ML and SageMaker training plans
Machine Learning Blog
This article explains how to secure GPU capacity for short-term ML workloads using EC2 Capacity Blocks for ML and SageMaker training plans.
- GPU demand exceeds supply; short-term capacity reservations address this challenge
- On-demand instances offer flexibility but uncertain availability; no cost advantage
- Spot instances reduce costs up to 90% but risk interruption; suitable for checkpointable workloads
- EC2 Capacity Blocks reserve GPU capacity for 1-182 days with 40-50% discount; supports up to 256 instances
- SageMaker training plans offer 70-75% discount for managed ML workloads; covers training, HyperPod, inference
- Decision framework: start with on-demand, use reserved capacity when timing certainty required
- Step-by-step guide provided for creating SageMaker training plans for inference endpoints
Choose reserved capacity options when workload timing is critical; otherwise start with on-demand for cost flexibility.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
May 18
2026
2026
Amazon SageMaker Studio now supports GPU capacity reservation through SageMaker Flexible Training Plans
Mar 24
2026
2026
Deploy SageMaker AI inference endpoints with set GPU capacity using training plans
Mar 17
2026
2026
SageMaker Training Plans now enables extending of existing capacity commitments without workload reconfiguration
Jan 11
2024
2024
Enhancing ML workflows with AWS ParallelCluster and Amazon EC2 Capacity Blocks for ML
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.