Run Spark 31% faster and optimize compute costs with Amazon S3 Express One Zone on Amazon EMR
Storage Blog
This article demonstrates how Amazon S3 Express One Zone improves Apache Spark performance on Amazon EMR through benchmarking with TPC-DS datasets.
- S3 Express One Zone reduced total benchmark runtime by 31% at 3TB scale and 47% at 10TB scale on an 8-node Graviton4 cluster
- 103 of 105 TPC-DS queries ran faster with no application code changes required
- Per-request latency dropped from 121ms to 7ms, a 17-fold reduction in read latency
- Overall cost per benchmark run decreased 36% due to shorter compute time and 85% lower S3 request pricing
- Recommended pattern: write data to both S3 Express One Zone and S3 Standard, use lifecycle rules to expire hot tier data after retention window
- Most I/O-heavy queries saw largest gains, with query q76 achieving 62.7% speedup
- Setup requires co-locating EMR cluster in same Availability Zone as directory bucket and disabling fast partition discovery
S3 Express One Zone delivers significant performance and cost benefits for I/O-bound analytics workloads, with benefits increasing as data volume grows.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Sep 1
2026
2026
Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent
Sep 3
2026
2026
Run Apache Spark up to 10x faster with DataPelago on Amazon EKS
Jul 29
2026
2026
Accelerate Spark on EMR Serverless with larger workers and shuffle-optimized disks
Jul 30
2025
2025
Optimize Amazon EMR runtime for Apache Spark with EMR S3A
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.