Home icon

Use Kubernetes Operators for new inference capabilities in Amazon SageMaker that reduce LLM deployment costs by 50% on average

Machine Learning Blog



This article discusses new inference capabilities in Amazon SageMaker that can help reduce large language model (LLM) deployment costs by 50% on average. It introduces the use of Kubernetes Operators for these new capabilities, allowing users to deploy and manage LLMs on SageMaker through Kubernetes.

Specifically, the article covers:

  • How AWS Controllers for Kubernetes (ACK) work, using Amazon S3 as an example
  • Key components of the new inference capabilities, including inference components
  • Solution overview for deploying Dolly v2 7B and FLAN-T5 XXL models on SageMaker using inference components and Kubernetes Operators
  • Prerequisites and steps for creating, updating, and deleting inference components via Kubernetes
  • Availability and pricing details for the new inference capabilities
  • Conclusion highlighting the benefits of using Kubernetes Operators for LLM deployment on SageMaker


Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Mar 18
2024
Optimize price-performance of LLM inference on NVIDIA GPUs using the Amazon SageMaker integration with NVIDIA NIM Microservices
Apr 8
2024
Boost inference performance for Mixtral and Llama 2 models with new Amazon SageMaker containers
Apr 23
2024
Accelerate ML workflows with Amazon SageMaker Studio Local Mode and Docker support
Apr 24
2024
Improve LLM performance with human and AI feedback on Amazon SageMaker for Amazon Engineering

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.