Home icon

Medidata’s journey to a modern lakehouse architecture on AWS

Big Data Blog



This article describes how Medidata, a clinical data platform company, modernized its data architecture from legacy batch ETL to a real-time lakehouse using AWS services and Apache Iceberg.

  • Replaced fragmented batch jobs with real-time Apache Flink streaming pipelines on Amazon EKS
  • Implemented Apache Iceberg tables backed by AWS Glue Data Catalog for unified data access
  • Reduced pipeline latency from days to minutes, achieving 99% performance improvement
  • Eliminated need for custom downstream pipelines through Iceberg interoperability
  • Reduced operational maintenance burden by five times with single data copy
  • Centralized security using AWS IAM instead of custom access control layers
  • Enabled point-in-time queries via Iceberg snapshots for data quality management
  • AWS Glue Iceberg optimizations handle compaction, snapshot retention, and orphan file deletion

Medidata's lakehouse architecture demonstrates how modern open-source technologies on AWS enable scalable, real-time clinical data platforms with improved performance, maintainability, and security.



Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Jan 12
2026
Navigating architectural choices for a lakehouse using Amazon SageMaker
Mar 3
2026
Building a modern lakehouse architecture: Yggdrasil Gaming’s journey from BigQuery to AWS
Dec 3
2024
AWS announces Amazon SageMaker Lakehouse
Jul 13
2026
Multi-cloud lakehouse architecture on AWS for Agentic AI, Part 1: Architecture and best practices

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.