Kafka 4.3 Adds Native Queues

What Data Engineers Should Know

IBM closed its $11.59B Confluent deal as Kafka 4.3 shipped production ready share groups. What native queues mean for data engineers.

Published: Proudly published by: Jeff Wahl10 min read

What this post covers

01

What Queues for Kafka Actually Changes: Native point to point delivery via share groups versus consumer groups

02

Jay Kreps Steps Down, Shaun Clowes Named CEO: Leadership transition after 12 years and new CEO's background

03

IBM Closes the $11.59 Billion Confluent Deal: Deal terms, closing timeline, and stated integration plans

04

Native Queues Versus SQS, RabbitMQ, and Bolt-On Queue Systems: When share groups replace a dedicated queue service in a pipeline

05

Kafka 4.3.0 Share Groups Reach Production Ready Status: Per message acknowledgment, redelivery, and cooperative semantics in 4.3.0

06

Kafka Streams Rebalance GA: What It Fixes: New rebalance protocol and reduced consumer downtime during scaling

07

Share Groups and Queues in Data Engineer Interview Prep: How new Kafka semantics show up in system design questions

08

What IBM Ownership Means for Kafka's Open Source Direction: Governance and roadmap questions for an Apache project under IBM

Kafka share groups are production ready. Apache Kafka 4.2.0 shipped on February 17, 2026 with Queues for Kafka (KIP-932) at general availability, which gives Kafka native point-to-point queue semantics for the first time.[1] Apache Kafka 4.3.0 followed on May 22, 2026 with 25 KIPs, more than 600 commits, and a set of share group configuration controls that operators had been waiting for.[2]

The ownership picture changed over the same period. IBM completed its acquisition of Confluent on March 17, 2026, paying $31 per share in cash. The deal is widely cited at $11.59 billion, and IBM's release puts the enterprise value at roughly $11 billion.[3] On August 13, 2026, co-founder Jay Kreps announced he was stepping down as CEO after almost 12 years and handing the role to Chief Product Officer Shaun Clowes.[4]

My read is that the queues feature matters more to your day job than the ownership change does. The ownership change still deserves attention, and it deserves a plan, which is different from panic. Below I cover what shipped, what it costs, what IBM actually bought, and how this will come up in your next Kafka interview loop.

Prepare for the interview
01 / Open invite
02min.

Know the patterns before the interviewer asks them.

a system design query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
PayPalInterview question
Solve a problem

What Queues for Kafka Actually Changes

For about a decade, the rule was simple: 1 partition goes to 1 consumer in a group. If you had 12 partitions and wanted 40 workers, you had 3 options. You could repartition the topic, bolt on some ugly external coordination, or pipe everything into RabbitMQ or SQS and run 2 messaging systems.

I've done that last option. I spent a year babysitting a RabbitMQ cluster that existed only because a job-dispatch workload didn't fit Kafka's partition model. It had its own on-call rotation, its own monitoring, and its own 3am failure modes. It was all overhead with no business value.

Share groups remove the partition ceiling. Multiple consumers in a share group can process records from the same partition at the same time, and each record is acknowledged individually instead of being tracked by a single committed offset.[5] The result is that you can have more consumers than partitions.

The cost is ordering. Consumer groups give you strict per-partition order because a single member owns each partition. Share groups hand records to whichever member asks next.

Consumer groupsShare groups
Assignment unitPartition, exclusive to 1 memberIndividual records, shared across members
Max useful consumersPartition countMore than partition count
Progress trackingCommitted offset per partitionPer-record acknowledgement and delivery state
OrderingStrict within a partitionNot guaranteed across members
DeliveryDepends on your commit strategyAt-least-once via acquisition locks
FitsSessionization, CDC, state machines, analytical pipelinesWork queues: rendering, invoicing, enrichment calls, backlog burn-down

The Acknowledgement Model

A share consumer resolves each record with one of 3 terminal acknowledgement types. ACCEPT means the record is done and won't be redelivered. RELEASE means a temporary failure, so the record goes back into the pool for another consumer. REJECT means a permanent failure, so the record is archived and never redelivered.[5] Kafka 4.2 added a 4th, non-terminal type, RENEW, which extends the acquisition lock on records that need more processing time.[1]

The broker keeps a delivery count for every record. Once a record hits share.delivery.count.limit, which defaults to 5, it is archived and stops being redelivered.[5] In practice, a single poison record can't crash-loop your workers indefinitely. Anyone who has watched a consumer group stall behind one malformed message knows why that matters.

Here is a minimal worker that uses explicit acknowledgement:

import java.time.Duration;
import java.util.List;
import java.util.Properties;
import org.apache.kafka.clients.consumer.AcknowledgeType;
import org.apache.kafka.clients.consumer.ConsumerRecord;
import org.apache.kafka.clients.consumer.ConsumerRecords;
import org.apache.kafka.clients.consumer.KafkaShareConsumer;
import org.apache.kafka.common.serialization.StringDeserializer;
public class InvoiceWorker {
static class TransientException extends RuntimeException {
TransientException(String msg) { super(msg); }
}
static void render(String invoiceJson) {
// call the PDF service; throw TransientException on timeouts
}
public static void main(String[] args) {
Properties props = new Properties();
props.put("bootstrap.servers", "localhost:9092");
props.put("group.id", "invoice-workers");
props.put("share.acknowledgement.mode", "explicit");
props.put("key.deserializer", StringDeserializer.class.getName());
props.put("value.deserializer", StringDeserializer.class.getName());
try (KafkaShareConsumer<String, String> consumer = new KafkaShareConsumer<>(props)) {
consumer.subscribe(List.of("invoices"));
while (true) {
ConsumerRecords<String, String> records = consumer.poll(Duration.ofMillis(500));
for (ConsumerRecord<String, String> record : records) {
try {
render(record.value());
consumer.acknowledge(record, AcknowledgeType.ACCEPT);
} catch (TransientException e) {
// back to the pool; delivery count increments
consumer.acknowledge(record, AcknowledgeType.RELEASE);
} catch (Exception e) {
// bad payload; archive it, never redeliver
consumer.acknowledge(record, AcknowledgeType.REJECT);
}
}
consumer.commitSync();
}
}
}
}

Delivery is at-least-once, so render() has to be safe to run twice. If you haven't built idempotent pipelines before, share groups will teach you quickly and at your expense.

Kafka Share Groups in 4.3: The Knobs That Make It Production Ready

Share groups took a while to reach this point. They were early access in Kafka 4.0, preview in 4.1, and GA in 4.2, which added RENEW, adaptive batching for share coordinators, and per-share-group lag metrics (KIP-1226).[1] Lag metrics are important. A queue you can't measure is one you'll learn about from someone else's Slack message.

Apache Kafka 4.3 shipped KIP-1240, which adds share-group-level configuration for acquisition lock duration, maximum in-flight records, and delivery attempt limits.[2] These used to be broker-wide defaults. Now you can tune each group, so the invoice workers and the thumbnail workers no longer have to share a retry policy because they happen to share a cluster.

That is the gap between a feature that works and a feature you can operate. Once you run multiple workloads on one cluster, per-group tuning is a requirement.

The CPU Bill

Per-record state costs something. In benchmarks published May 22, 2026, Jack Vanlightly found that share groups use measurably more CPU than consumer groups because of per-record accounting and state management. Throughput was broadly comparable and could exceed consumer groups in high-fanout scenarios.[6]

His results come from one set of benchmarks. Your hardware, record sizes, and acknowledgement patterns will change the numbers, so run your own test before migrating anything high-volume. The direction of the result makes sense, though. The broker now tracks the state of every record, and that bookkeeping uses CPU.

Key takeaway: Share groups let you consolidate a separate queue system onto Kafka and scale workers past your partition count. You pay for it with ordering and some broker CPU. For work queues that trade is usually good. For anything where event order is part of correctness, keep your consumer groups.

Should Share Groups Replace SQS or RabbitMQ?

Sometimes. Ask yourself 2 questions.

  • Is the data already in Kafka? If events land in Kafka and then get copied into SQS so workers can pick them up, you're running a glue service whose only job is to move data between queues. Share groups can remove it.
  • Do you need push-style, per-message latency? Kafka consumers still pull and batch. If a workload needs single-digit-millisecond dispatch per message, a push-based broker may still be the better tool. Measure before assuming either way.

The economics usually decide it. A second messaging system comes with a second on-call rotation, a second set of dashboards, and a second skill set to hire for. Engineer time costs far more than broker CPU. If share groups can absorb the workload, turning off the extra system is often the biggest saving available.

Kafka Streams Rebalance: GA in 4.2, Filled Out in 4.3

The release timeline is less clear than the headlines suggest, so here it is directly. The 4.2.0 announcement said the Kafka Streams server-side rebalance protocol (KIP-1071) reached GA with a limited feature set: sticky assignment only, offline-only migration, and no warm-up tasks or rack-aware task placement.[1] The 4.3.0 announcement also presents the protocol as GA, highlighting a sticky task assignor that minimizes task movement and a new StreamsGroupDescribe RPC for streams-specific metadata.[2] My interpretation is GA in 4.2 and more complete in 4.3. The 2 announcements frame it differently, and I'd rather point that out than hide it.

The fix itself is significant. Under the old model, the client-side rebalance acted as a global synchronization barrier, and stateful Streams apps could stall while every instance agreed on a new assignment. KIP-1071 moves assignment computation to the broker, and the sticky assignor keeps tasks where they already are whenever it can.[2] Kafka 4.3 also includes KIP-1035, which stores changelog offsets inside each state store instead of in a separate checkpoint file.[2]

Opting in is a single config line in your Streams properties:

application.id=sessionizer
bootstrap.servers=localhost:9092
group.protocol=streams

Changing that line is easy. Migrating is not. Moving an existing app from the classic protocol has been offline-only, so plan for a maintenance window and read the upgrade notes for your exact patch version before starting. If your app relies on warm-up tasks to keep failover fast, check whether that is supported yet before you switch.[1]

The Confluent IBM Acquisition, by the Numbers

The facts on the Confluent IBM acquisition come straight from the close announcement: $31 per share, all cash, completed March 17, 2026. Confluent serves more than 6,500 enterprises, including 40% of the Fortune 500.[3] IBM's stated plan is to integrate Confluent's streaming platform with its own data and integration stack, including watsonx.data, positioning real-time data as the foundation for enterprise AI and agents.[3]

One point gets muddled constantly on Reddit. IBM bought Confluent, the company, including its commercial platform, Confluent Cloud, and its enterprise products. IBM did not buy Apache Kafka. Kafka is still an Apache Software Foundation project under the Apache 2.0 license, and its own PMC controls releases and technical direction.

That separation is real, and it is also imperfect. Confluent employs many of the people who write Kafka, and a large share of the KIPs discussed here came from Confluent engineers. Governance can stay independent while funding and engineering hours shift. That is where I'd focus attention. A license change is unlikely. A drift in which KIPs get staffed is more plausible.

Analysts Are Slowing the Store Down

> We run an e-commerce marketplace where the analytics team queries the production database directly, and that load is degrading the live application. Move analytics onto its own warehouse by reading the database's change log instead of querying the live system, while a merchant-facing dashboard still shows each seller their new orders within fifteen minutes on a path of its own. A small fraction of orders arrive with broken merchant references or totals that do not add up, so those have to be held back and caught before they reach the reporting tables.

+ Source
+ Transform
+ Storage
+ Quality
+ Consumer
+ Queue
Bronze
Silver
Gold
Custom
Pipeline Architecture
Sketch the architecture.

Click or drag a node from the toolbar above. Right-click the canvas for the full menu.

Drag from a node's right port to another node's left port to wire data flow.

Jay Kreps Steps Down as CEO; Shaun Clowes Takes Over

On August 13, 2026, Kreps posted that he was stepping back after almost 12 years and handing the CEO role to Shaun Clowes. "There is an ambitious roadmap for what's next," he wrote, "and I'll watch from the sidelines with a lot of pride and a little FOMO as the team achieves that."[4]

Clowes joined Confluent as Chief Product Officer in December 2022 from MuleSoft, where he was CPO and led a team of more than 250 product, design, and content specialists across automation, integration, and API management.[7]

Industry coverage described the departure as a blow to IBM so soon after closing the deal. That is a fair reading of the optics. Founders leaving within a year of an acquisition is also extremely common, and I've been through enough acquisitions to call it the default outcome. The data that would actually tell us something hasn't arrived yet: the Confluent-contributed KIP count over the next 4 releases, and whether IBM bundles Confluent features into larger IBM product suites.

Here is my speculation, labeled as speculation. A CEO who spent years running integration and API management products fits well with IBM's integration plans. Expect Confluent's commercial roadmap to lean toward connectors, governance, and hybrid deployment. Whether that helps you depends on whether you use IBM's stack.

Kafka Interview Questions Data Engineers Should Expect Now

I've sat on enough hiring panels to know how new features get into interviews. First, someone on the panel reads the release notes. About 6 months later, the question shows up in 30% of loops. Get ahead of that cycle.

Expect these in the Kafka portion of a data engineering system design round:

  • "You have 12 partitions and need 50 workers. What do you do?" The old answer was to repartition. The current answer is a share group, provided the work doesn't depend on order. Say that condition out loud.
  • "When would you still pick a consumer group?" Anything where order is part of correctness, such as CDC replay, sessionization, or per-account state machines. Also any throughput-heavy case where the CPU overhead Vanlightly measured actually matters.
  • "A record keeps failing. What happens?" Walk through RELEASE versus REJECT, the delivery count limit, and why the handler has to be idempotent under at-least-once delivery.
  • "Why do Streams rebalances cause lag spikes, and what changed?" Explain eager, then cooperative sticky, then broker-side assignment under KIP-1071. Mention the current limitations, because that's what separates people who have run it from people who have only read about it.
  • "Kafka or a managed queue?" This is now a consolidation question. Weigh operational cost against push latency and ordering requirements.

The underlying concept is the same one it has always been: stream versus queue, ordered log versus independent work items. The API changed, the concept didn't. If you understand why ordering and elastic parallelism conflict, you can answer any version of this question with any tool. Drill the fundamentals with the Kafka interview questions set, and if you still default to streaming when batch would work, reread batch vs streaming trade-offs. Most of y'all still don't need streaming, and saying so clearly in an interview is a senior-level signal.

What Data Engineers Should Do Next

Based on what has shipped and what is still unclear, here's what I'd do:

  • Inventory your side queues. Find every SQS or RabbitMQ deployment that only exists because Kafka couldn't scale workers past the partition count. Those are your share group candidates.
  • Benchmark before migrating. Use your own record sizes and acknowledgement patterns. Vanlightly's CPU findings tell you where to look, not what your numbers will be.[6]
  • Target 4.3 for share groups. The per-group tuning from KIP-1240 is what makes multi-tenant clusters manageable.[2]
  • Treat Streams protocol migration as a project. Offline migration, a scheduled window, and patch-version upgrade notes.
  • Separate "Kafka" from "Confluent" in your risk planning. Apache Kafka's roadmap is public and governed by the PMC. Confluent Cloud's roadmap is now set inside IBM. If you depend on Confluent-only connectors or schema tooling, write down your exit path now while you have time.

For interview prep, take these patterns into a mock interview and practice explaining the stream-versus-queue decision under time pressure. It's a different skill from knowing it.

Over the next 12 months, watch 2 things: how many KIPs from Confluent engineers land in 4.4 and 4.5, and whether IBM starts bundling Confluent into larger product suites. If both stay steady, the leadership change is just a normal org chart update. Either way, you'll probably still be debugging consumer lag at 2am. That part doesn't change.

References

  1. Apache Kafka, "Apache Kafka 4.2.0 Release Announcement," February 17, 2026. kafka.apache.org
  2. Apache Kafka, "Apache Kafka 4.3.0 Release Announcement," May 22, 2026. kafka.apache.org
  3. IBM Newsroom, "IBM Completes Acquisition of Confluent, Making Real Time Data the Engine of Enterprise AI and Agents," March 17, 2026. newsroom.ibm.com
  4. Jay Kreps, post on X announcing CEO transition, August 13, 2026. x.com
  5. Confluent Documentation, "Share Consumers for Confluent Platform." docs.confluent.io
  6. Jack Vanlightly, "Benchmarking Apache Kafka Consumer Groups vs Share Groups (overhead test)," May 22, 2026. jack-vanlightly.com
  7. Confluent Investor Relations, "Shaun Clowes Appointed Chief Product Officer of Confluent," December 2022. investors.confluent.io

Kafka share groupsQueues for KafkaConfluent IBM acquisitionJay Kreps CEOApache Kafka 4.3Kafka Streams rebalance
02 / Why practice

Try the actual problems

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes