Optimize GPU workloads on Amazon EKS and ROSA with IBM Turbonomic
IBM and Red Hat Blog
This article explains how IBM Turbonomic optimizes GPU-powered generative AI inference workloads on Amazon EKS and ROSA by continuously analyzing application demand against GPU supply and generating optimization actions.
- Turbonomic integrates Kubeturbo and Prometurbo components to monitor GPU utilization, memory, and application-level metrics like response time and request queuing
- Generates four types of optimization actions: horizontal scaling, container resizing, pod rebalancing, and GPU placement decisions
- Correlates GPU signals with application SLOs to determine appropriate responses—scaling out when performance degrades, consolidating when utilization is low
- Supports NVIDIA GPU sharing mechanisms including time-slicing, Multi-Instance GPU (MIG), and Dynamic Resource Allocation (DRA)
- Step-by-step implementation includes connecting clusters, adding Prometheus monitoring, validating resource discovery, and enabling optional automation policies
- Supports MIG partition optimization on A100, H100, and H200 GPUs for hardware-isolated workload consolidation
Turbonomic provides an application-aware optimization layer that maintains SLOs while improving GPU utilization and reducing infrastructure costs for inference services.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2024
2025
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.