Amazon SageMaker HyperPod now supports disaggregated prefill and decode
News
Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD), an inference optimization that separates LLM inference phases onto dedicated GPU pools.
- Separates compute-bound prefill and memory-bandwidth-bound decode onto different GPU pools to eliminate resource contention
- Delivers consistent per-token latency, higher throughput at strict latency SLOs, and independent scaling of prefill/decode capacity
- Intelligent router automatically directs long-context requests through disaggregated path while routing short prompts directly to decoder
- Enabled via pdSpec section in InferenceEndpointConfig custom resource on HyperPod Inference Operator
- Composable with existing KV cache offloading and intelligent routing features
- Available on EKS orchestrator with EFA-capable instance types across all AWS Regions supporting SageMaker HyperPod
DPD enables production LLM deployments to achieve better latency predictability and throughput efficiency under mixed traffic patterns.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Jul 10
2026
2026
Disaggregated prefill and decode for LLM inference on SageMaker HyperPod
Jul 2
2026
2026
Amazon SageMaker HyperPod now supports AMI versioning and auto-patching
May 20
2026
2026
Amazon SageMaker HyperPod now supports data capture for inference workloads
Jun 1
2026
2026
Amazon SageMaker HyperPod now supports EFA-only network interfaces
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.