Home icon

Build a centralized observability platform for Apache Spark on Amazon EMR on EKS using external Spark History Server

Big Data Blog



This AWS Big Data Blog article details how to build a centralized observability platform for Apache Spark on Amazon EMR on EKS using an external Spark History Server (SHS), providing a unified monitoring solution across multiple clusters.

  • Solution enables collecting Spark events from multiple EMR on EKS clusters into a central S3 bucket
  • Deploys Spark History Server on a dedicated Amazon EKS cluster
  • Secures access using AWS Load Balancer Controller, AWS Private CA, Route 53, and AWS Client VPN
  • Integrates DataFlint for enhanced performance monitoring and insights
  • Provides a single, secure interface to monitor and troubleshoot Spark applications

The solution addresses the complexity of monitoring Spark applications across diverse workloads by creating a centralized, secure, and comprehensive observability platform.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

May 1
2025
Build end-to-end Apache Spark pipelines with Amazon MWAA, Batch Processing Gateway, and Amazon EMR on EKS clusters
Jun 12
2025
Implement observability for Amazon EKS workloads using the Instana Amazon EKS Add-on
May 29
2025
Optimizing data lakes with Amazon S3 Tables and Apache Spark on Amazon EKS
Jun 5
2025
Using AWS Glue Data Catalog views with Apache Spark in EMR Serverless and Glue 5.0

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.