How Common Crawl and AWS Open Data built the foundation for the AI revolution
Public Sector Blog
This article describes how Common Crawl and AWS Open Data built the foundation for the AI revolution by providing free access to web-scale training data for large language models.
- Common Crawl has hosted 300 billion web pages on AWS S3 since 2012 at no cost through the AWS Open Data Sponsorship Program
- 64% of major LLMs published 2019–2023 used filtered Common Crawl data for training, cited in 13,000+ research papers
- Infrastructure evolved to include S3, CloudFront, Athena, EMR, Lambda, and SageMaker for global access and processing
- Enables researchers, builders, and organizations worldwide to access foundational AI training data without building crawling infrastructure
- Web Languages Project and CommonLID benchmark address multilingual AI representation across 109+ languages
Open data infrastructure has become essential to AI development, enabling researchers globally to access the same foundational datasets powering frontier AI systems.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
Aug 24
2026
2026
Democratizing institutional knowledge: Building an AI-powered knowledge management system with AWS
Sep 24
2026
2026
Announcing the New AWS Reimagine Report on AI
Aug 18
2026
2026
Powering agentic AI with real-time streaming data on AWS
Sep 11
2026
2026
Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.