Serverless analytics pipelines using the Apache Spark engine in Amazon Athena
Big Data Blog
This article demonstrates how to build serverless analytics pipelines using the Apache Spark engine in Amazon Athena with Spark Connect support, eliminating traditional cluster management overhead.
- Pattern A: Interactive analysis with Jupyter notebooks for exploratory data analysis and feature engineering
- Pattern B: Local development with VS Code for building production-ready Spark applications
- Pattern C: Scheduled pipelines with dbt and Apache Airflow for version-controlled transformations
- Spark Connect provides secure, authenticated remote connectivity with automatic scaling up to 60 workers
- Sessions launch in seconds with pay-per-use pricing and 20-minute idle timeout by default
- Live Spark UI and History Server support debugging; AWS Lake Formation integration secures Glue Data Catalog access
- Session-level cost attribution enables granular chargeback and budgeting in AWS Cost Explorer
Athena with Apache Spark transforms Spark workload operations by providing near-instant serverless compute, allowing data teams to focus on analytics rather than infrastructure management.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2024
2025
2026
2024
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.