Medidata’s journey to a modern lakehouse architecture on AWS
Big Data Blog
This article describes how Medidata, a clinical data platform company, modernized its data architecture from legacy batch ETL to a real-time lakehouse using AWS services and Apache Iceberg.
- Replaced fragmented batch jobs with real-time Apache Flink streaming pipelines on Amazon EKS
- Implemented Apache Iceberg tables backed by AWS Glue Data Catalog for unified data access
- Reduced pipeline latency from days to minutes, achieving 99% performance improvement
- Eliminated need for custom downstream pipelines through Iceberg interoperability
- Reduced operational maintenance burden by five times with single data copy
- Centralized security using AWS IAM instead of custom access control layers
- Enabled point-in-time queries via Iceberg snapshots for data quality management
- AWS Glue Iceberg optimizations handle compaction, snapshot retention, and orphan file deletion
Medidata's lakehouse architecture demonstrates how modern open-source technologies on AWS enable scalable, real-time clinical data platforms with improved performance, maintainability, and security.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2026
2024
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.