Snowflake Now Speaks Kafka

What Data Engineers Need to Know

Snowflake Datastream lets existing Kafka producers connect with a config change and land data as Iceberg tables in seconds. What Data Engineers should evaluate now.

Published: Proudly published by: Jeff Wahl9 min read

What this post covers

01

The $6 Billion AWS Deal and What It Signals: 5-year AWS infrastructure commitment, what the bet means for platform direction

02

How Data Lands: Iceberg or Native Tables in Seconds: Seconds-to-query latency, Iceberg vs native table landing, data freshness guarantees

03

What Snowflake Datastream Actually Is: Kafka protocol support, config-only producer migration, fully managed service

04

Datastream vs a Self-Managed Kafka Cluster: Operational burden, cost tradeoffs, latency, control, when each fits

05

What Kafka-Compatible Means in Practice: Producer compatibility scope, consumer-side behavior, protocol fidelity limits

06

Governance by Default: Horizon Catalog on Every Topic: RBAC, lineage, classification, masking inherited automatically from Horizon Catalog

07

Streaming Architecture in Data Engineering Interviews in 2026: How Datastream shifts system design questions and expected answers this year

At Snowflake Summit 2026, Snowflake announced Snowflake Datastream: a fully managed streaming service that speaks the Apache Kafka wire protocol. Existing Kafka producers connect with a config change. No SDK swap, no code rewrite. Data lands as native Snowflake or Apache Iceberg tables, queryable within seconds of arrival. Every topic automatically inherits Horizon Catalog governance: RBAC, classification, lineage, masking.[1]

I've been saying streaming is overrated for years, and I still believe that for 90% of teams. But for the teams that genuinely need sub-minute latency into their warehouse, this announcement changes the operational math. Here's what matters and what doesn't.

Prepare for the interview
01 / Open invite
02min.

Know the patterns before the interviewer asks them.

a system design query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
PayPalInterview question
Solve a problem

What Snowflake Datastream Actually Is

Datastream is a Kafka-compatible streaming service native to Snowflake. It speaks the full Kafka wire protocol, so your existing producers (librdkafka, the Java client, whatever you're running) point at a Datastream endpoint instead of your Kafka broker, and they work.[1] A config change.

Snowflake is direct about what's underneath. Streaming analyst Kai Waehner noted: "Snowflake is refreshingly direct about Datastream: it is compatible with the Kafka wire protocol but does not use Kafka under the hood. The technology is pure Snowflake."[2] Under the covers, Datastream uses a stateless architecture writing directly to blob storage, with metadata in a highly available in-memory store. That's fundamentally different from Kafka's local broker state model, and it means Snowflake can scale and govern the streaming layer the same way it governs everything else on the platform.

Snowflake's EVP of Product Christian Kleinerman framed the motivation at Summit: customers wanted "a Kafka-type of streaming solution, but in the Snowflake way," with managed infrastructure, automatic scaling, and governance baked in from the start. That's a revealing quote. Snowflake isn't trying to replace Kafka the event backbone; they're replacing Kafka the ingestion pipe.

The service entered private preview in June 2026, with early customers getting access through the summer.[3] Pricing and GA timelines haven't been published yet.

What Kafka-Compatible Streaming Means in Practice

The phrase "Kafka-compatible" is doing heavy lifting, and it's worth understanding the edges before your team makes architecture decisions based on a marketing slide.

Producer side: compatibility is real. If you're running a Kafka producer that serializes events and ships them to a broker endpoint, Datastream accepts that traffic with a config swap. The wire protocol is the contract, and Datastream honors it.[1] Exactly-once delivery semantics are supported through offset token tracking; consumers commit progress and replay from the last committed offset on recovery, preventing duplicates and data loss.

Consumer side: the model is different. Datastream topics land as Snowflake tables. You query them with SQL, not with a Kafka consumer client pulling from a topic. If your architecture depends on consumer groups, multi-consumer fan-out, or Kafka Streams applications reading from the same topic independently, Datastream doesn't serve those patterns. It was never designed to.

Waehner framed this well: "The Kafka API has become the de facto standard for moving events around, the same way the Amazon S3 API became the standard for object storage. However, using the Kafka API for ingestion into an analytics platform represents a different architectural approach than running Kafka itself."[2]

Datastream is Kafka-compatible for ingestion into Snowflake. If Snowflake is your only destination, the distinction between Datastream and Kafka is academic. If Snowflake is 1 of 8 consumers, Datastream doesn't replace the broker; it replaces the connector.

For anyone building Kafka fundamentals for interviews, the event-driven architecture concepts (consumer groups, partitioning, exactly-once semantics, backpressure) still matter regardless of whether a managed service handles the plumbing.

How Data Lands: Seconds, Not Minutes

Data from Datastream topics lands as either native Snowflake tables or Apache Iceberg tables, queryable within seconds of ingestion.[1] The architecture uses zero-copy streaming with sub-second latency, separating storage and compute to minimize data movement.[3]

For comparison, Snowflake's existing Snowpipe Streaming handles up to 20 GB/s throughput per table with data queryable in approximately 5 seconds.[4] Datastream claims to improve on that latency, though detailed production benchmarks aren't public yet given the private preview status.

The Iceberg landing option is the more interesting architectural choice. Iceberg tables let other engines (Spark, Trino, Flink) read the same data without routing through Snowflake's compute. That's a hedge against lock-in for teams running multi-engine architectures. Native Snowflake tables give you tighter query performance inside Snowflake. The trade-off is interoperability versus speed, and you make it per-topic based on who reads the data downstream. For teams building toward a lakehouse architecture, the Iceberg path preserves optionality without giving up the managed ingestion layer.

This fits a broader industry pattern. Confluent Tableflow, StreamNative Ursa, Databricks Zerobus, and Fluss are all converging on the same idea: turning the topic into a table without an intermediate ETL step. The streaming layer and the analytical layer are collapsing into a single surface. That convergence is the trend worth tracking, regardless of which vendor you're on.

Governance Without the Side Quest

This is where Datastream gets genuinely compelling, and it's the part most coverage skips.

Every Datastream topic automatically inherits full Horizon Catalog governance at ingestion time. RBAC, tag-based classification, column-level lineage, and dynamic masking policies apply the moment data arrives, without manual configuration.[1] You don't set up a separate governance layer for streaming data. You don't bolt on a schema registry and hope somebody configured the ACLs. The governance is the platform.

If you've worked on a team where the warehouse had proper access controls and the Kafka topics were the Wild West, you understand why this matters. (Everyone has worked on that team.) Streaming data is historically the governance blind spot. It moves fast, it's high volume, and nobody wants to slow down the pipeline to add masking policies. So the policies never get applied, and then compliance asks questions nobody can answer.

Snowflake also announced Intent-Driven Governance alongside Datastream, which converts plain-language policy templates into active Horizon Catalog rules. Combined with Datastream's automatic inheritance, a policy set once propagates to every new topic without further human intervention. That's governance as code, applied at ingestion, covering every stream. For teams under SOC 2 or GDPR obligations, that removes an entire category of "we'll get to it later" backlogs.

The consolidation has a cost. Pulling governance, billing, and ingestion inside a single vendor's perimeter increases lock-in. That's a real trade-off. But for teams already running warehouse, transformation, and BI on Snowflake, adding governed streaming to the same stack is a simpler total-cost-of-ownership calculation than most vendor consolidation plays. It's an economics argument, and the economics are hard to argue with when you price out the alternative.

The Math on Self-Managed Kafka

Architecture decisions get made on spreadsheets, not whiteboards. Here are the numbers.

A production Kafka deployment typically requires 0.5 to 2 full-time engineers for patching, upgrades, incident response, and 24/7 monitoring. That's before infrastructure costs. A 3-node production cluster generates $14,000 to $24,000 per month in cross-AZ replication traffic alone, often accounting for over 50% of total infrastructure costs.[5] Teams routinely underestimate total Kafka operational burden by 3x to 5x at budget time. I've watched this happen firsthand; the initial estimate is always "we'll run it on 3 brokers and tune it once," and 18 months later there are 2 engineers whose calendars are 40% Kafka.

Add it up: 6 figures a year on infrastructure and headcount to move data from a topic to a table. If Snowflake is the primary destination for that data, Datastream collapses that cost into your existing Snowflake bill.

But if Kafka serves 8 consumers and Snowflake is consumer number 4, tearing out the cluster makes no sense. The economics only work when Snowflake is the primary (or sole) sink. I see teams about to make bad decisions in both directions: Snowflake-only shops keeping Kafka alive out of inertia, and multi-consumer shops ripping out Kafka because Datastream sounds easier.

I've been consistent on this: most companies don't need streaming at all. A daily batch job running in 20 minutes at $5 beats a streaming pipeline costing $500/day with a dedicated on-call engineer. But for the subset that genuinely needs near-real-time analytics in Snowflake, Datastream removes the largest chunk of operational tax.

Analysts Are Slowing the Store Down

> Analytics queries against the production database are slowing the live application. Move analytics onto its own warehouse fed from the database's change log, while a merchant dashboard shows new orders within fifteen minutes on a path of its own.

+ Source
+ Transform
+ Storage
+ Quality
+ Consumer
+ Queue
Bronze
Silver
Gold
Custom
Pipeline Architecture
Sketch the architecture.

Click or drag a node from the toolbar above. Right-click the canvas for the full menu.

Drag from a node's right port to another node's left port to wire data flow.

The $6 Billion AWS Signal

3 weeks before Summit, Snowflake announced a 5-year, $6 billion AWS infrastructure commitment: annual minimums ranging from $900 million to $1.25 billion through March 2031.[6] For context, Snowflake's AWS commitment was $1.2 billion at their 2020 IPO, $2.5 billion in 2023, and $6 billion now; a 2.4x increase over the previous deal.[7]

The market responded accordingly. Shares surged 37% in after-hours trading on Q1 FY27 results showing product revenue up 34% year-over-year to $1.33 billion.[8] The deal focuses on Graviton processors and GPU-accelerated compute for AI workloads, with over 13,600 accounts already using Snowflake's AI capabilities.

Why does this matter for Datastream? Infrastructure commitments this large precede product roadmap changes. Snowflake is betting that its platform becomes the default substrate for real-time analytics and AI workloads, and Datastream is the ingestion layer that feeds them. You don't spend $6 billion on cloud infrastructure if the product roadmap is "better batch queries." The streaming and AI layers are load-bearing for the next 5 years of Snowflake's strategy, and the AWS deal tells you exactly how much they're willing to spend to prove it.

Snowflake reaffirmed their multi-cloud commitment across AWS, Azure, and GCP. But $6 billion buys prioritization. Teams on AWS plus Snowflake should expect the deepest integrations and the fastest feature velocity.

What This Changes for Data Engineering Interviews

If you're preparing system design answers for Snowflake-centric roles in late 2026, the vocabulary has shifted.

The old answer to "Design a real-time ingestion pipeline" involved standing up Kafka, configuring a schema registry, building a consumer that writes to the warehouse, bolting on monitoring, and discussing partition strategies. That answer still holds for multi-consumer architectures.

The new answer, for single-destination Snowflake workloads, is shorter: point producers at Datastream, query the table. The complexity shifts from infrastructure plumbing to governance design. "How would you tag and mask PII in a streaming topic?" is more relevant than "How do you tune Kafka broker configs?" for teams on this stack. But here's the thing interviewers actually want to hear: you know both answers and can articulate why you're choosing one over the other. That's the difference between a mid-level and a staff-level response.

The meta-skill hasn't changed: know when to pick which architecture and cost-justify it. Clarify the business latency requirement first. If the answer is "daily dashboards," you don't need streaming. If the answer is "fraud detection under 500ms," you need a real event backbone. If the answer is "near-real-time analytics already in Snowflake," Datastream has the smallest operational surface. The strongest candidates in 2026 aren't the ones who know the most tools; they're the ones who can reason about total cost of ownership under time pressure.

For Snowflake-specific prep, Datastream is now part of the product surface you should know. For pipeline architecture knowledge more broadly, the distinction between "streaming as infrastructure" and "streaming as ingestion" is the concept that transfers across platforms. Learn the concept; the product names are syntax.

What to Watch

Datastream is in private preview. You can't run it in production today. But you can prepare.

  • Audit your Kafka topology. Count consumers per topic. If Snowflake is the only destination for most topics, Datastream is worth evaluating the moment it hits GA.
  • Know the boundary. Kafka-compatible ingestion is not Kafka. Multi-consumer fan-out, Kafka Streams, event sourcing: those still need a broker.
  • Watch the Iceberg angle. Landing as Iceberg tables gives you a multi-engine read path that native Snowflake tables don't. If you're choosing between formats, pick based on who needs to read the data, not on what sounds cooler.
  • Think governance first. Automatic Horizon Catalog inheritance is the feature with the widest impact. If your streaming data has weaker governance than your warehouse data, that gap is closable without a separate project.

The tools change every 18 months. The problems don't change. Schema drift, late-arriving data, upstream teams breaking contracts without telling you: those are eternal. Datastream is a new tool for old problems. Evaluate it on the operational math, not the hype cycle.

References

  1. Snowflake, "Snowflake Datastream: Kafka-Native Streaming in Snowflake." snowflake.com
  2. Kai Waehner, "Why Databricks and Snowflake Speak the Kafka Protocol: Ingestion vs. Architecture," June 22, 2026. kai-waehner.de
  3. select.dev, "Snowflake Summit 2026: Product Announcement Recap," June 2026. select.dev
  4. Snowflake, "Snowflake Snowpipe Streaming High-Performance Architecture." snowflake.com
  5. AutoMQ, "Kafka Cross-AZ Hidden Cost," 2026. automq.com
  6. Snowflake, "Snowflake Expands AWS Collaboration with $6B Commitment to Accelerate Enterprise Agentic AI Adoption," May 27, 2026. snowflake.com
  7. GeekWire, "Snowflake commits $6B to Amazon Web Services over 5 years in latest AI infrastructure deal," May 28, 2026. geekwire.com
  8. Yahoo Finance, "Snowflake Explodes 37% on $6 Billion Amazon Deal as CEO Calls Q1 an AI 'Inflection Point,'" May 28, 2026. finance.yahoo.com
Snowflake DatastreamKafka-compatible streamingSnowflake Summit 2026managed streaming data engineeringdata engineer streaming architecture
02 / Why practice

Try the actual problems

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes