Apache Iceberg v3 Is GA

What Data Engineers Need to Know

Apache Iceberg v3 is GA on Snowflake, Databricks, and AWS. Deletion vectors and row lineage are the 2 changes data engineers should act on first.

Published: Proudly published by: Jeff Wahl8 min read

What this post covers

01

Row Lineage and the Incremental Processing Payoff: Row-level change tracking and what it cuts from CDC and incremental refresh costs

02

4 More New Capabilities in the v3 Spec: Default column values, geometry types, nanosecond timestamps, and multi-argument partition transforms

03

VARIANT Type: Native Semi-Structured Data Without Workarounds: New VARIANT column type replacing JSON-in-string and external schema patterns

04

3 Platforms, 1 Spec, Simultaneously: Why This Is New: Snowflake, Databricks, and AWS S3 Tables all GA on Iceberg v3 at once

05

Deletion Vectors: What Changes for Update-Heavy Pipelines: How deletion vectors reduce write amplification versus copy-on-write rewrites

06

Iceberg 1.11.0: Spark 4.1 and Flink 2.1 as Default Build Targets: Library upgrade compatibility and what it means for local pipeline tooling

07

What Data Engineers Should Know for Interviews and Day-to-Day Work: Table format questions now tested in interviews; v3 semantics as a production baseline skill

I've spent the last 3 years watching data engineers argue about table formats like it's a religious war. Delta vs Iceberg vs Hudi; each camp armed with benchmarks that conveniently prove their format wins. That argument just got a lot quieter.

Apache Iceberg v3 reached general availability on Snowflake on May 7, 2026.[1] Databricks shipped it to GA in June 2026.[2] AWS S3 Tables has been rolling out v3 features since November 2025, with VARIANT support landing July 28, 2026.[3] All 3 major cloud platforms converged on the same lakehouse table format spec within weeks of each other. The new spec includes deletion vectors, row lineage, a VARIANT type, default column values, geography and geometry types, and nanosecond timestamps.[1] If you're building on a lakehouse architecture, Iceberg v3 is the production baseline everywhere that matters.

Prepare for the interview
01 / Open invite
02min.

Know the patterns before the interviewer asks them.

a system design query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
PayPalInterview question
Solve a problem

3 Platforms, 1 Spec: Why This Convergence Is New

This has never happened with a table format release. Iceberg v2 landed on different platforms months apart. Delta Lake was Databricks-first for years before other engines caught up. Hudi's adoption timeline was even more fragmented. You always had to check platform support before designing around a feature.

With v3, the answer is yes across the board. Snowflake's GA dropped May 7 with the full feature set.[1] Databricks entered public preview on April 24, 2026 (Runtime 18.0+) and hit full GA in June, covering managed Iceberg tables, foreign Iceberg tables, and UniForm-enabled managed tables.[2] AWS S3 Tables added deletion vectors and row lineage across Spark, AWS Glue, and SageMaker notebooks in November 2025, then completed the v3 feature set with VARIANT in July 2026.[3]

What does simultaneous GA mean in practice? You can design a pipeline around v3 semantics and run it on any of these platforms without feature gating. Your deletion vectors work the same on Snowflake and Databricks. Your VARIANT columns are queryable everywhere. The portability promise that Iceberg has always made actually holds at the spec level now, across the vendors that employ most data engineers.

Snowflake went further: on June 1, 2026, Snowflake-managed storage for Iceberg tables reached GA, eliminating external S3 volumes, IAM configuration, and manual compaction.[4] Automatic compaction, garbage collection, encryption, Time Travel, and Fail-Safe are handled natively.[4] That's Snowflake betting most teams don't want to manage object storage alongside an open table format. Whether that bet ages well depends on how much you trust a single vendor with your storage layer.

Deletion Vectors: The Biggest Win for Update-Heavy Pipelines

If you've run a MERGE on a large Iceberg v2 table and watched it rewrite hundreds of gigabytes of Parquet files to update a few thousand rows, deletion vectors are the fix.

v2 used copy-on-write for all DML: any UPDATE, DELETE, or MERGE rewrote entire data files to produce new versions without the deleted rows. A 1GB file with 1 changed row meant rewriting the full gigabyte. Write amplification on update-heavy workloads was brutal.

v3 flips the model. Instead of rewriting files, the engine stores a binary bitmap per data file marking deleted rows. At read time, the engine checks the bitmap for O(1) lookup and skips the marked rows. AWS EMR benchmarks show up to 10x faster DML operations compared to v2's copy-on-write approach.[5]

This matters most for incremental upserts, SCD Type 2 maintenance, and any pattern where you're touching a subset of rows in a large table. Operations that used to require careful batching and off-peak scheduling to manage I/O now cost a fraction of the compute.

Deletion vectors make Iceberg v3 the default answer for update-heavy lakehouse pipelines. The write amplification tax that made v2 painful for CDC and SCD workloads is effectively gone.

-- v3 MERGE: deletion vectors handle the old row versions
-- Only a bitmap + new data files are written; source Parquet files stay untouched
MERGE INTO orders AS target
USING staging_orders AS source
ON target.order_id = source.order_id
WHEN MATCHED AND source.status != target.status THEN
UPDATE SET
status = source.status,
updated_at = current_timestamp()
WHEN NOT MATCHED THEN
INSERT (order_id, customer_id, status, created_at, updated_at)
VALUES (source.order_id, source.customer_id, source.status,
source.created_at, current_timestamp());

The SQL is identical to what you'd write on v2. The execution plan isn't. Under v2, this MERGE rewrites every data file containing a matched row. Under v3 with deletion vectors, it writes a small bitmap marking old row versions and appends new data files only for updated and inserted rows. Same semantics; dramatically less I/O.

Row Lineage and the Incremental Processing Payoff

v3 adds native row lineage fields to every row: a row ID and a last-modified sequence number baked directly into the format.[6]

Before v3, building CDC pipelines on Iceberg required external tooling. You'd track changes through snapshot diffs, maintain watermark tables, or run full-table comparisons to figure out what changed. Every team rolled their own version, and every version had edge cases around late-arriving data, compaction rewriting row order, and concurrent writes. I've debugged enough of these home-rolled CDC setups to know they all break in the same 3 ways; they just break on different schedules.

With native row IDs and sequence numbers, the table itself tells you which rows changed and when. Your incremental pipeline reads the lineage fields instead of diffing snapshots. Less code to maintain, fewer edge cases, and a standardized contract across engines.

For anyone building idempotent pipelines with incremental refresh patterns, row lineage eliminates an entire category of custom infrastructure. The sequence number gives you a total ordering of mutations per row; the row ID gives you stable identity across compaction. Those 2 properties are what every CDC framework was trying to approximate from the outside.

The VARIANT Type: Semi-Structured Data in Apache Iceberg v3

Every data engineer has a pipeline that ingests JSON. Every one of those pipelines hits the same decision point: parse into typed columns at ingest (and deal with schema evolution when upstream adds a field at 4am on a Friday), or dump into a STRING column and parse at query time (and accept terrible performance forever).

The VARIANT type is the middle ground that actually works. VARIANT stores JSON in a compact binary format, supporting field-level queries without deserializing the entire document.[7] Frequently accessed fields get "shredded" into typed Parquet columns automatically, so the query engine reads them as native typed data while the rest of the document stays in its binary representation.[7]

You get the flexibility of schemaless ingest with performance that approaches typed columns for your hot fields. Storage cost is lower than text JSON. Query cost drops because the engine skips string parsing entirely.

-- Querying a VARIANT column
-- Shredded fields (user_id, action) read as typed Parquet columns under the hood
SELECT
event_id,
event_data:user_id::INT AS user_id,
event_data:action::STRING AS action,
event_data:metadata:device_type::STRING AS device_type
FROM events
WHERE event_data:action::STRING = 'purchase'
AND event_data:metadata:platform::STRING = 'ios';

AWS S3 Tables shipped VARIANT support on July 28, 2026, completing v3 coverage across all 3 platforms.[3] If you've been maintaining custom JSON parsing logic in your ingestion layer, VARIANT is the migration worth prioritizing.

Analysts Are Slowing the Store Down

> Analytics queries against the production database are slowing the live application. Move analytics onto its own warehouse fed from the database's change log, while a merchant dashboard shows new orders within fifteen minutes on a path of its own.

+ Source
+ Transform
+ Storage
+ Quality
+ Consumer
+ Queue
Bronze
Silver
Gold
Custom
Pipeline Architecture
Sketch the architecture.

Click or drag a node from the toolbar above. Right-click the canvas for the full menu.

Drag from a node's right port to another node's left port to wire data flow.

Default Values, Geometry Types, and Nanosecond Timestamps

Deletion vectors, row lineage, and VARIANT get the headlines, but v3 ships additional features that eliminate common workarounds:[1]

  • Default column values. Columns can declare defaults in the schema. Adding a new column to a table with billions of rows doesn't require a backfill job. Reads of older data files return the default for the new column automatically. Less compute, less coordination, faster schema evolution.
  • Geography and geometry types. Native spatial data types in the format spec. Geospatial columns survive round-trips across engines without casting to WKB strings and back.
  • Nanosecond timestamps. v2 capped at microsecond precision. v3 adds nanoseconds, which matters for high-frequency trading data, IoT sensor streams, and any workload where microsecond granularity loses information.

None of these are individually revolutionary. Together, they eliminate workarounds that data engineers have been maintaining in application code for years. Default values alone will save hours of migration planning for any team doing regular schema evolution on production tables.

Apache Iceberg 1.11.0: Spark 4.1 and Flink 2.1 as Default Targets

While the cloud platforms shipped v3 support in their managed services, the open-source library got its own major release. Apache Iceberg 1.11.0 dropped in May 2026, built from over 1,000 commits by 200+ contributors.[8]

The practical headline: Spark 4.1 and Flink 2.1 are the default build targets.[8] If you're running self-managed Spark or Flink, this is your upgrade signal. The REST catalog now plans scans server-side, shifting metadata work from query engines to the catalog service.[8] That means less memory pressure on your executors and better query planning for large tables with thousands of partitions.

The release also includes built-in envelope encryption with Google KMS support and a new partition statistics scan API for optimizer access to table shape.[8] If you haven't pinned your Iceberg library version recently, 1.11.0 is the one to pin to.

What Data Engineers Should Know for Interviews

Table format questions are showing up in interviews at the platform companies and the startups building on top of them. The questions have moved past "what is Iceberg?" and into "when would you choose deletion vectors over copy-on-write?" and "how does row lineage change your CDC architecture?" I've been on both sides of the table for too many loops to count; the questions evolve, but the underlying skill stays the same. Understanding how your data is physically organized and why it matters.

Here's what to focus on:

  • Deletion vectors vs copy-on-write. Know the tradeoff cold. Copy-on-write gives you clean reads (no bitmap check at scan time) but punishes write-heavy workloads with full file rewrites. Deletion vectors optimize for writes at the cost of a bitmap lookup on reads. For most production workloads with mixed read/write patterns, deletion vectors win. This is a pipeline design question now.
  • Row lineage for CDC. Explain how native row IDs and sequence numbers eliminate external CDC tooling. Interviewers want to hear that you understand the operational cost of maintaining custom change tracking infrastructure.
  • VARIANT vs parsed schemas vs STRING columns. Know why VARIANT is the better middle ground for semi-structured data. The shredding optimization separates "I read the docs" from "I've used this in production": hot fields get typed Parquet columns, everything else stays binary.
  • Platform defaults. Snowflake requires explicit v3 table creation; Databricks defaults to v3 on new tables.[2] Knowing these differences signals production awareness beyond spec knowledge.

Concepts transfer across tools; tool knowledge doesn't transfer across concepts. Deletion vectors, row lineage, and VARIANT aren't Iceberg trivia. They're the latest answers to problems DEs have always faced: how to update data efficiently, how to track changes, and how to handle semi-structured data without losing your mind. If you're prepping for data engineering interviews, add v3 semantics to your study plan alongside SQL and Python fundamentals. The format is new; the problems are eternal.

The table format wars produced a clear outcome: convergence on a single open spec across every major platform. For the first time, learning one format's semantics gives you coverage on Snowflake, Databricks, and AWS simultaneously. That's rare in this industry. Build on it.

References

  1. Snowflake Documentation, "May 7, 2026: Support for Apache Iceberg Version 3 (General availability)." docs.snowflake.com
  2. Databricks Blog, "Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance." databricks.com
  3. AWS What's New, "Amazon S3 Tables now support the Variant data type for Apache Iceberg V3," July 28, 2026. aws.amazon.com
  4. Snowflake Documentation, "Jun 1, 2026: Snowflake storage for Apache Iceberg tables (General availability)." docs.snowflake.com
  5. Atlan, "Apache Iceberg v3: New Features and Snowflake 2026 Guide," citing AWS EMR benchmarks. atlan.com
  6. Snowflake Builders Blog on Medium, "Iceberg V3 & Snowflake: Unlocking Bidirectional Change Data Capture." medium.com
  7. Dremio, "The VARIANT Type: How to Store JSON Without the Pain." dremio.com
  8. Google Open Source Blog, "Announcing Apache Iceberg 1.11.0," May 2026. opensource.googleblog.com

Apache Iceberg v3deletion vectorsIceberg GA 2026lakehouse table formatdata engineer Iceberg
02 / Why practice

Try the actual problems

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes