Home icon

Amazon SageMaker AI inference now supports G7 instances

News



This article announces support for G7 instances in Amazon SageMaker AI inference, powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs.

  • Up to 4.6x AI inference performance improvement compared to G6 instances
  • 32 GB GPU memory per GPU with 5th Generation Tensor Cores
  • Up to 700 Gbps EFA-enabled networking (7x faster than G6)
  • Up to 7.6 TB local NVMe SSD storage for large model deployment
  • Ideal for 7B–30B parameter models, image/video generation, and multi-model endpoints
  • Deploy via SageMaker console, API, or SDK with ml.g7.xlarge through ml.g7.48xlarge
  • Available in US East (N. Virginia, Ohio) and US West (Oregon)

G7 instances enable cost-effective production inference for large generative AI models without over-provisioning or model quantization.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Apr 20
2026
Accelerate Generative AI Inference on Amazon SageMaker AI with G7e Instances
Nov 22
2024
Amazon SageMaker Inference now supports G6e instances
Dec 11
2024
Amazon SageMaker AI announces availability of P5e and G6e instances for Inference
May 4
2026
Amazon SageMaker AI Now Supports Capacity-Aware Inference with Automatic Instance Fallback

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.