Running Red Hat AI on OpenShift with AWS Neuron
IBM and Red Hat Blog
This article demonstrates how to run Red Hat AI on OpenShift using AWS Neuron chips for LLM inference, featuring the matured AWS Neuron Operator v1.1.4.
- AWS Neuron Operator reached v1.1.4 with stable v1.x API and Neuron SDK 2.28 support
- Supports rolling driver upgrades across nodes without cluster downtime
- Includes node metrics telemetry and automated CI/CD workflows
- Deploy via OpenShift GitOps (Argo CD) for production lifecycle management
- Run Llama 3.1 8B inference using Red Hat AI Inference vLLM Neuron image
- Supports both raw Kubernetes Deployment and KServe-based model serving approaches
- Production optimizations: PVC compilation cache, EBS snapshots, S3 model storage
- Requires ROSA cluster with inf2/trn1 instances and cluster-admin privileges
The post provides step-by-step deployment instructions using GitOps, testing procedures, and production optimization strategies for running LLM inference on AWS Neuron hardware within OpenShift environments.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2024
2024
2025
2024
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.