The Databricks Delta Lake Iceberg debate used to be a real architectural decision. You picked a format, committed your storage layer, and lived with the consequences. If you needed to switch, you rewrote petabytes of Parquet files. At Data+AI Summit 2026 (June 15-18), Databricks announced that Delta Lake and Apache Iceberg now share the same underlying Parquet data files, with format selection reduced to a metadata configuration change.[1] Simultaneously, Databricks Lakebase hit general availability, putting managed Postgres on lakehouse storage governed by Unity Catalog.[2] These aren't incremental feature releases. They collapse 2 of the biggest architectural decisions data engineers have been making for the last 5 years into configuration switches.
Delta Lake Meets Iceberg on Databricks
Data Engineer Guide
Databricks now runs Delta Lake and Iceberg on shared Parquet files and shipped Lakebase GA at Data+AI Summit 2026. What it means for Data Engineers.
What this post covers
What the Delta vs. Iceberg Question Looks Like Now: How shared storage changes the format selection calculus
Format Switching Without a Full Data Rewrite: Switching formats as metadata choice, not storage copy
Delta Sharing and the Iceberg REST Catalog: Sharing live Delta data to any Iceberg REST Catalog recipient
Databricks Lakebase GA: Postgres on Lakehouse Storage: Serverless Postgres with compute decoupled on the lakehouse
What Shared Parquet Storage Actually Means: Delta and Iceberg tables reading identical physical Parquet files
Unity Catalog Governing OLTP and OLAP Together: Single governance layer across transactional and analytical workloads
Data Engineer Interview: Lakehouse Architecture in 2026: How these announcements shift system design interview answers
Know the patterns before the interviewer asks them.
What Shared Parquet Storage Actually Means
Delta Lake and Iceberg have always been metadata layers on top of Parquet. The data files themselves were already columnar Parquet. The difference was how each format tracked transactions, schema evolution, and file-level statistics. The problem was that those metadata layers were incompatible, so running both formats meant maintaining 2 copies of the same data.
UniForm (Universal Format) eliminates the duplication. When you write a Delta table with UniForm enabled, Databricks asynchronously generates Iceberg metadata alongside the Delta transaction log, pointing at the same Parquet files. No data rewrite. No second copy. One set of Parquet files, 2 metadata layers.[3]
The write overhead is negligible. Even for petabyte-scale tables, the metadata is a tiny fraction of total data size, and UniForm incrementally generates metadata scoped only to changes since the previous commit.[4] This matters because the economics argument against dual-format support was always about storage cost. At 2 cents per GB per month, duplicating a 100TB dataset costs $2,000/month for no analytical value. UniForm kills that cost.
The convergence goes deeper. Iceberg v3 standardized deletion vectors using the same Roaring Bitmap implementation as Delta Lake. Deletion vectors mark rows as deleted in separate bitmap files rather than rewriting the underlying Parquet data, delivering up to 10x faster DELETE, UPDATE, and MERGE operations.[1] Because the binary encodings are identical across both formats, deletion vectors work seamlessly through UniForm.
"It doesn't matter if it's in Delta or Iceberg. It's actually the same now." , Ali Ghodsi, Databricks CEO, Data+AI Summit 2026[5]
And the roadmap makes this permanent. Databricks announced that Iceberg v4 and Delta 5.0 will converge on a unified metadata tree by Q4 2026, producing a single on-disk structure that both Delta and Iceberg clients read and write directly with no translation layer.[6] The format war is resolving into a format merger.
Format Switching Without a Full Data Rewrite
Before UniForm, switching from Delta to Iceberg (or the reverse) meant reprocessing your entire dataset. I've seen teams spend weeks planning migrations that were functionally just rewriting the same Parquet files with different metadata. Engineering time that could've gone toward, you know, building something useful.
Now the calculus is different. Enabling UniForm on a new Delta table is a table property. Enabling it on an existing table requires a full rewrite only if the table uses deletion vectors (which are the default in newer Delta versions).[7] After that, every Iceberg-compatible client (Snowflake, BigQuery, Trino, Flink, DuckDB) reads the same Parquet files through Iceberg metadata. No dual ingestion pipelines. No replication.
The practical effect: format selection is now workload-driven, not tribal. Delta Lake remains stronger for Spark-native teams that need liquid clustering and Change Data Feed for CDC patterns. Iceberg offers wider engine support across Snowflake, BigQuery, Trino, Flink, and DuckDB, making it the safer default for multi-engine deployments.[1] But the key word is "default." With shared storage, you're not locked in. Organizations can deploy both formats within the same Unity Catalog, using Delta for latency-sensitive operational analytics and Iceberg for cross-platform data sharing.
If you're studying lakehouse architectures for interviews, understand this shift. The old answer to "Delta or Iceberg?" was a 10-minute architectural analysis. The new answer starts with "both, on the same files" and pivots to governance and engine ecosystem decisions.
Delta Sharing Meets the Iceberg REST Catalog
Shared storage solves internal format interoperability. Delta Sharing solves external data exchange. In January 2026, Databricks announced first-class Iceberg format support in Delta Sharing, enabling data recipients to consume Delta Shares in any Iceberg-compatible client.[8]
The numbers suggest this isn't a niche feature. Delta Sharing has achieved more than 300% year-on-year usage growth for 2 consecutive years and serves more than 16,000 data recipients across clouds, platforms, and regions. 40% of those connections flow through non-Databricks platforms: Apache Spark, pandas, PowerBI, Tableau, Excel.[8]
Data providers can now share live Iceberg tables from external catalogs (AWS Glue, Snowflake Horizon, Hive Metastore) to any client supporting the Apache Iceberg REST Catalog API. Unity Catalog handles credential vending, issuing scoped, short-lived tokens to external engines without requiring static credentials.[9] This replaces the pattern where teams exchanged long-lived S3 keys or service account JSON files to share data across organizations. If you've ever managed a spreadsheet of cross-account IAM roles for data sharing, you understand why credential vending matters.
The Linux Foundation also announced the OpenSharing Project on June 10, 2026, providing a community-governed protocol supporting Delta Sharing, Iceberg, and Parquet formats for exchanging data and AI assets across organizational boundaries.[10] This standardization push signals that data sharing protocols are heading toward the same kind of convergence as table formats.
Databricks Lakebase: Postgres on Lakehouse Storage
Lakebase is the announcement that'll change how you think about OLTP versus OLAP architecture. Databricks acquired Neon for $1 billion in May 2025 to build this, and the result is a fully managed, serverless Postgres service storing operational data directly on lakehouse storage governed by Unity Catalog.[2]
Lakebase hit GA on AWS on February 3, 2026, and on Azure on March 3, 2026.[2] Adoption grew at more than 2x the rate of Databricks' data warehousing product, with thousands of companies running production workloads. The platform now handles 12 million database launches per day.[11]
The technical architecture separates compute from storage. Postgres compute scales elastically (including scale-to-zero after a configurable idle period, default 5 minutes, with reactivation in hundreds of milliseconds). Storage capacity reaches 8TB per instance on PostgreSQL 17.[12] Instant database branching uses copy-on-write semantics, so you can create zero-copy branches of production data in seconds without duplicating underlying storage.[13]
At Data+AI Summit 2026, Databricks announced LTAP (Lake Transactional/Analytical Processing), the architectural pattern that unifies OLTP and OLAP on a single copy of data. Lakebase's Postgres engine writes to its WAL and buffer cache, then converts rows to columnar format before landing in object storage, making the data immediately readable through Delta and Iceberg table formats.[11]
I'll be honest about the trade-off here. LTAP creates 2 consistency domains rather than a unified transaction domain. There's measurable lag between transactional and analytical reads. For real-time inventory dashboards that need sub-second consistency with the OLTP system, that lag matters. For the 90% of analytical workloads where data that's a few seconds old is fine, it doesn't. Know which category your use case falls into before you pitch LTAP in a system design interview.
Unity Catalog: Governing OLTP and OLAP Together
The governance story is where these announcements compound. Unity Catalog now governs Delta tables, Iceberg tables, Lakebase Postgres databases, and shared datasets through a single identity, permissions, and audit model.[11] One catalog for your transactional application data and your analytical warehouse. One set of access controls. One audit trail.
Catalog Commits shift transaction coordination from the file system to Unity Catalog, enabling multi-table transactions across Delta and Iceberg tables while maintaining ACID guarantees. Requires Databricks Runtime 16.4 and above.[14]
There's a governance gap worth knowing about. Unity Catalog controls analytical access through Databricks compute, but direct Postgres connections to Lakebase bypass Unity Catalog's access controls because query execution occurs entirely within PostgreSQL.[15] If your security model depends on Unity Catalog for row-level security on Lakebase data, you need to ensure all access routes through Databricks compute, not direct Postgres connections. This is the kind of detail that separates a surface-level understanding from a production-grade one.
For teams managing medallion architectures with bronze, silver, and gold layers, the governance consolidation is significant. Your raw ingestion layer (bronze), your cleaned and conformed layer (silver), your business-ready aggregates (gold), and now your application database can all live under the same catalog with consistent access policies.
Lakehouse Architecture in Data Engineer Interviews
These announcements shift what a strong system design answer looks like. If you're preparing for Databricks interviews or any senior data engineering role, here's what changed.
Format Selection Questions
The old answer: "We chose Delta because we're a Spark shop" or "We chose Iceberg for multi-engine support." The new answer: "We write Delta with UniForm enabled, giving us Iceberg compatibility for downstream consumers on Snowflake and Trino, while our Spark jobs read through the Delta transaction log for liquid clustering and CDC support." Format selection is now about engine ecosystem and workload characteristics, not storage commitment.
OLTP + OLAP Architecture Questions
The old answer: "We run Postgres for the application, CDC into Kafka, land in S3, process with Spark into Delta tables." The new answer acknowledges that Lakebase collapses the first 3 steps. Strong answers still distinguish between use cases. Real-time inventory with sub-second consistency requirements? You might still want a separated architecture. Analytics on application data where a few seconds of lag is fine? LTAP eliminates the CDC pipeline entirely. Knowing when to collapse versus separate is the signal interviewers look for in system design rounds.
Governance and Multi-Engine Questions
With format interoperability solved, interview focus shifts from "which format wins?" to "how do we govern this?" Understanding REST Catalog standards, credential vending, and Unity Catalog's governance boundaries (including the Lakebase direct-connection gap) demonstrates depth. If you're asked about cross-organizational data sharing, reference Delta Sharing's Iceberg REST Catalog support and the OpenSharing protocol rather than proposing bespoke replication pipelines.
The Concepts Still Transfer
Here's what hasn't changed: you still need to understand why columnar storage matters for analytical queries, why transaction logs exist, how schema evolution works, and what ACID guarantees actually mean at the storage layer. The tools are converging; the concepts underneath them are the same concepts they've always been. If you understand data modeling fundamentals and storage-layer trade-offs, adapting to UniForm or Lakebase is a weekend of reading docs. If you only know the API surface of one tool, you're starting over every 18 months.
What to Watch Next
The Iceberg v4 and Delta 5.0 unified metadata tree, targeted for Q4 2026, is the announcement that makes all of this permanent.[6] If that ships on schedule, the Delta versus Iceberg question becomes roughly equivalent to asking "MySQL or MariaDB?" The underlying storage and metadata are the same; the client libraries and ecosystem integrations are the differentiators.
Lakehouse RT, also announced at summit, promises sub-100ms query latency on Delta Lake and Iceberg tables through a new compute engine called Reyden.[1] If that delivers, it collapses yet another tier: the real-time serving layer that teams currently build on Redis or DynamoDB for low-latency reads.
The pattern is clear. Databricks is systematically eliminating the architectural tiers that data engineers have spent years building and maintaining: separate OLTP and OLAP systems, separate formats for separate engines, separate serving layers for different latency requirements. Each eliminated tier is one less pipeline to build, monitor, and debug at 2am.
The tools change. The problems (governance, consistency, latency trade-offs, cost optimization) don't change. Learn the concepts; the syntax is the easy part.
References
- Databricks, "Advancing the Lakehouse with Apache Iceberg v3 on Databricks," June 2026. databricks.com
- Databricks, "Azure Databricks Lakebase is Generally Available," March 2026. databricks.com
- Databricks, "Delta Lake Universal Format (UniForm) for Iceberg compatibility, now in GA," June 2024. databricks.com
- Databricks, "Delta UniForm: a universal format for lakehouse interoperability." databricks.com
- Medium, "Data + AI Summit 2026: What Databricks Actually Announced," June 2026. medium.com
- Databricks, "Format Co-Evolution: How Iceberg v4 and Delta 5.0 Share a Unified Metadata," Data+AI Summit 2026 session, June 2026. databricks.com
- Databricks, "Read Delta Lake tables with Iceberg clients using UniForm." docs.databricks.com
- Databricks, "Announcing first-class support of Iceberg format in Databricks Delta Sharing," January 23, 2026. databricks.com
- Databricks, "Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance." databricks.com
- Linux Foundation, "Linux Foundation Announces OpenSharing Project to Standardize AI Asset and Data Exchange," June 10, 2026. linuxfoundation.org
- Databricks, "Databricks Launches LTAP: The First Lake Transactional/Analytical Processing Architecture," June 16, 2026. databricks.com
- Databricks, "Beyond Provisioning: Developer's Guide to Databricks Lakebase Autoscaling." docs.databricks.com
- Databricks, "Database Branching in Postgres: Git-Style Workflows with Databricks Lakebase." databricks.com
- Databricks, "Catalog commits," official documentation. docs.databricks.com
- Databricks, "Register a Lakebase database in Unity Catalog." docs.databricks.com
Try the actual problems
- 01
Reading a solution is not the same as writing one
Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you
- 02
76% of hiring managers reject on the coding task, not the resume
From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice
- 03
The AI round turns on how you drive the tool, not whether you use it
Specifying the change, reviewing the diff, catching the wrong edit, iterating to green. Doing it on a timer is what turns 'I use Copilot' into a defensible workflow
Related interview prep
senior data engineer interview guide
Senior Data Engineer interview process, scope-of-impact framing, technical leadership signals.
FAANG data engineer interview questions
Real questions from Meta, Amazon, Apple, Netflix, and Google Data Engineer loops, with answers.
system design round prep guide
Pipeline architecture, exactly-once semantics, and the framing that gets you to L5.
The weekly data challenge
Dirty, production-shaped data, scored blind each week.