Best practices for scaling large consumer groups on Amazon MSK
Big Data Blog
This article provides best practices for scaling large consumer groups on Amazon MSK by managing metadata record sizes that can exceed the default 1 MB limit during rebalances.
- Consumer group metadata grows with member count, topic names, and client IDs; use formula: member_count × (3 × topic_name_bytes + 2 × client_id_bytes + ~200) to estimate size
- Increase max.message.bytes on __consumer_offsets topic and set replica.fetch.max.bytes at cluster level to handle larger metadata records
- Split large consumer groups into multiple smaller groups to reduce per-group metadata size proportionally
- Right-size partition counts and configure auto scaling bounds to prevent unbounded consumer group growth
- Optimize naming conventions (shorter group names and client IDs) to reduce per-member overhead
- Monitor HeapMemoryAfterGC, UnderReplicatedPartitions, and rebalance metrics; set CloudWatch alarms at 60% and 80% heap usage
- Use broker sizing guidance: kafka.m5.xlarge for 500-1,000 members; kafka.m5.2xlarge for 1,000-2,000 members; kafka.m5.4xlarge for 2,000+ members
- KIP-848 in Apache Kafka 4.0 reduces metadata size by computing assignments server-side instead of per-member
Combining increased message limits with complementary strategies like group splitting and capacity planning enables operating consumer groups at required scale while maintaining broker health.
The AWS News Feed is currently looking for gold sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.
Related articles
2026
2024
2026
2026
The AWS News Feed is currently looking for silver sponsors. If you want to support the AWS community and reach a large audience of AWS professionals, consider sponsoring the AWS News Feed.