Build a data lakehouse in a hybrid Environment using Amazon EMR Serverless, Apache DolphinScheduler, and TiDB
Big Data Blog
This article provides a comprehensive guide to building a data lakehouse in a hybrid environment using Amazon EMR Serverless, Apache DolphinScheduler, and TiDB. The solution addresses data synchronization and job orchestration challenges for enterprises with data-sensitive applications.
- Key components include:
- Amazon EMR Serverless for serverless data processing
- Apache DolphinScheduler for job orchestration
- TiDB as an on-premises or cloud data warehouse
- Amazon S3 for data storage
- AWS Glue Data Catalog for metadata management
- Data synchronization methods:
- Using TiDB Dumpling for historical data export
- TiDB CDC connector for incremental data sync
- EMR Serverless Spark jobs for data transfer
- Key benefits:
- Simplified data operations
- Flexible hybrid cloud architecture
- Optimized resource utilization
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Mar 4
2025
2025
Build a data lake for streaming data with Amazon S3 Tables and Amazon Data Firehose
Jul 28
2025
2025
Accelerate your data quality journey for lakehouse architecture with Amazon SageMaker, Apache Iceberg on AWS, Amazon S3 tables, and AWS Glue Data Quality
Dec 3
2024
2024
Amazon DynamoDB zero-ETL integration with Amazon SageMaker Lakehouse
Jul 14
2025
2025
Build real-time data lakes with Snowflake and Amazon S3 Tables
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.