Home icon

Migrate large HPC datasets from the edge to the cloud then synchronize continuously

Storage Blog



This article describes a solution for migrating large High-Performance Computing (HPC) datasets from on-premises locations with limited bandwidth to the AWS cloud, and then continuously synchronizing the data. It covers the following key points:

Specifically, the article covers:

  • Using AWS Snowball Edge Storage Optimized devices for initial bulk data transfer
  • Using the snow-transfer-tool to efficiently transfer data to Snowball devices
  • Using AWS DataSync for post-migration synchronization of updated data and metadata
  • Mounting migrated data on Amazon FSx for Lustre for low-latency access to HPC workloads
  • Detailed steps for ordering Snowball devices, transferring data, configuring DataSync, and setting up FSx for Lustre
  • Cleaning up resources after migration to avoid future charges


Go to article

The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.

Related articles

Aug 1
2024
Unlock scalability, cost-efficiency, and faster insights with large-scale data migration to Amazon Redshift
Jul 30
2024
HPC Ops: DevOps for HPC workloads in the cloud
Jul 25
2024
Migrate workloads from AWS Data Pipeline
Jul 2
2024
Improve HPC workloads on AWS for environmental sustainability

The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.