Kafka Share Groups in 4.3: The Knobs That Make It Production Ready
Share groups took a while to reach this point. They were early access in Kafka 4.0, preview in 4.1, and GA in 4.2, which added RENEW, adaptive batching for share coordinators, and per-share-group lag metrics (KIP-1226).[1] Lag metrics are important. A queue you can't measure is one you'll learn about from someone else's Slack message.
Apache Kafka 4.3 shipped KIP-1240, which adds share-group-level configuration for acquisition lock duration, maximum in-flight records, and delivery attempt limits.[2] These used to be broker-wide defaults. Now you can tune each group, so the invoice workers and the thumbnail workers no longer have to share a retry policy because they happen to share a cluster.
That is the gap between a feature that works and a feature you can operate. Once you run multiple workloads on one cluster, per-group tuning is a requirement.
The CPU Bill
Per-record state costs something. In benchmarks published May 22, 2026, Jack Vanlightly found that share groups use measurably more CPU than consumer groups because of per-record accounting and state management. Throughput was broadly comparable and could exceed consumer groups in high-fanout scenarios.[6]
His results come from one set of benchmarks. Your hardware, record sizes, and acknowledgement patterns will change the numbers, so run your own test before migrating anything high-volume. The direction of the result makes sense, though. The broker now tracks the state of every record, and that bookkeeping uses CPU.
Key takeaway: Share groups let you consolidate a separate queue system onto Kafka and scale workers past your partition count. You pay for it with ordering and some broker CPU. For work queues that trade is usually good. For anything where event order is part of correctness, keep your consumer groups.
Should Share Groups Replace SQS or RabbitMQ?
Sometimes. Ask yourself 2 questions.
- Is the data already in Kafka? If events land in Kafka and then get copied into SQS so workers can pick them up, you're running a glue service whose only job is to move data between queues. Share groups can remove it.
- Do you need push-style, per-message latency? Kafka consumers still pull and batch. If a workload needs single-digit-millisecond dispatch per message, a push-based broker may still be the better tool. Measure before assuming either way.
The economics usually decide it. A second messaging system comes with a second on-call rotation, a second set of dashboards, and a second skill set to hire for. Engineer time costs far more than broker CPU. If share groups can absorb the workload, turning off the extra system is often the biggest saving available.