Home icon

Serverless analytics pipelines using the Apache Spark engine in Amazon Athena

Big Data Blog



This article demonstrates how to build serverless analytics pipelines using the Apache Spark engine in Amazon Athena with Spark Connect support, eliminating traditional cluster management overhead.

  • Pattern A: Interactive analysis with Jupyter notebooks for exploratory data analysis and feature engineering
  • Pattern B: Local development with VS Code for building production-ready Spark applications
  • Pattern C: Scheduled pipelines with dbt and Apache Airflow for version-controlled transformations
  • Spark Connect provides secure, authenticated remote connectivity with automatic scaling up to 60 workers
  • Sessions launch in seconds with pay-per-use pricing and 20-minute idle timeout by default
  • Live Spark UI and History Server support debugging; AWS Lake Formation integration secures Glue Data Catalog access
  • Session-level cost attribution enables granular chargeback and budgeting in AWS Cost Explorer

Athena with Apache Spark transforms Spark workload operations by providing near-instant serverless compute, allowing data teams to focus on analytics rather than infrastructure management.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Sep 4
2024
Use Apache Spark on Amazon EMR Serverless directly from Amazon Sagemaker Studio
Nov 21
2025
Amazon EMR Serverless now supports Apache Spark 4.0.1 (preview)
Jan 26
2026
Apache Spark 4.0.1 preview now available on Amazon EMR Serverless
Dec 10
2024
Run Apache Spark Structured Streaming jobs at scale on Amazon EMR Serverless

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.