How Ramp runs GPU AI workloads at scale with ECS Managed Instances
Containers Blog
This article describes how Ramp, a finance automation platform, migrated GPU-powered AI inference workloads from manually managed ECS on EC2 to Amazon ECS Managed Instances, reducing operational complexity.
- Ramp's ML models made 26 million decisions across $10 billion in spend in October 2025, requiring GPU-accelerated infrastructure
- Original ECS on EC2 setup required manual management of Auto Scaling groups, launch templates, AMI patching, and monitoring scripts
- ECS Managed Instances eliminated fleet management by having AWS handle patching, scaling, and instance lifecycle
- Ramp uses separate capacity providers per GPU type (L40S and A10G) with instance_requirements instead of pinned instance types
- Multi-AZ deployment prevents capacity constraints; built-in autoscaler optimizes GPU utilization through bin-packing
- GPU workloads now follow same deployment patterns as Fargate services, improving consistency and onboarding speed
- Migrated 50-60 GPU instances across Bore and Embeddings workloads with no production downtime
ECS Managed Instances brought GPU workloads into operational parity with Ramp's broader ECS infrastructure, freeing engineers from manual fleet management while maintaining access to full EC2 capabilities.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.