Apache Airflow 3.3

State Store and Java/Go Task SDK

Airflow 3.3 adds a task and asset state store and a Java and Go Task SDK. What changed since July 2026 and what it means for data engineers.

Published: Proudly published by: Jeff Wahl9 min read

What this post covers

01

Expanded Asset Partitioning in 3.3: What changed for asset-aware scheduling since Airflow 3.0

02

Multi-Language Task SDK: Java and Go Task Logic: Writing task logic outside Python for the first time in Airflow

03

Pluggable Retry Policies Replace Fixed Backoff Defaults: Custom retry logic now configurable per task or DAG

04

What Shipped in Airflow 3.3, 3.3.1, and 3.3.2: Release timeline, scope, and point-release fixes since July 2026

05

The New State Store for Tasks and Assets: What a first-class state store changes for pipeline and asset design

06

Why Polyglot Pipelines Matter for Data Engineering Teams: How non-Python task logic affects teams with mixed-language stacks

07

How 3.3 Builds on the Airflow 3.0 Asset-Centric Rewrite: Placing 3.3's changes in context of Airflow 3's redesign

08

What Airflow 3.3 Means for Data Engineer Orchestration Skills: How new Airflow capabilities show up in system design interviews

Apache Airflow 3.3 shipped on July 6, 2026, and it's the most consequential Airflow release since the 3.0 rewrite. It adds 4 things: a first-class state store for tasks and assets (AIP-103), a Task SDK that runs task logic written in Java and Go (AIP-108), pluggable retry policies (AIP-105), and 3 new asset partition mappers.[1] Point releases followed on August 12 (3.3.1) and September 17 (3.3.2), mostly security and stability fixes.[3][4]

The headline for working data engineers is the state store. It's the first structural change to how Airflow tracks task and asset state since 3.0 went asset-centric. The Java and Go SDK gets more attention, but it's explicitly experimental.[2] Adoption context matters too: Astronomer's State of Apache Airflow 2026 report puts Airflow 3 at 26% of all users, with 84% planning to upgrade.[7] Most of you are reading about features you can't use yet.

Here's what shipped, what it replaces, and which parts deserve your time.

Prepare for the interview
01 / Open invite
02min.

Know the patterns before the interviewer asks them.

a system design query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
PayPalInterview question
Solve a problem

What Shipped in Apache Airflow 3.3, 3.3.1, and 3.3.2

ReleaseDateWhat matters
3.3.0July 6, 2026Task and asset state store, Java/Go Task SDK (experimental), pluggable retry policies, RollupMapper / FanOutMapper / FixedKeyMapper partition mappers, roughly 42x faster mapped task cleanup
3.3.1August 12, 2026Security fixes: team-scoped secret lookups returning unrelated teams' secrets, masking of sensitive Variables stored as JSON lists, secrets recorded unmasked in audit logs on bulk updates
3.3.2September 17, 2026Task SDK IPC fixes, Calendar view hangs on high-frequency cron, Calendar view computing runs in UTC instead of the DAG's timezone

Sources for the table: the 3.3.0 announcement and release notes, plus the 3.3.1 and 3.3.2 release notes.[1][3][4]

If you're on 3.3.0 and run multi-team deployments, the 3.3.1 secret-lookup fix alone is reason to upgrade. A secrets backend handing Team A's credentials to Team B is the kind of bug you'd rather hear about from a release note than an auditor.

The Airflow State Store: Durable Task State Without the Hacks

The Airflow project describes the problem better than I can: "Previously, if a task needed to remember something across retries or runs (a cursor, a checkpoint, a high-water mark) you reached for XComs, an external store, or a clever Variable hack. Airflow 3.3 makes durable task state a first-class concept."[1]

Mechanically, tasks get a task_state_store accessor and assets get asset_state_store. Both hold arbitrary key-value state that survives retries and run boundaries. State lives in the metadata database by default, or in a custom worker-side backend. It supports per-key retention with garbage collection and an optional clear_on_success flag.[2]

The resumable job pattern

I have personally been paged for this exact bug more than once. A task submits a long-running Spark job, polls it, and the worker dies mid-poll. Airflow retries the task. The task submits the job again. Now 2 copies are writing to the same table and somebody downstream is asking why revenue doubled overnight.

Every team I've been on solved this with a homegrown workaround. One used a Variable keyed on DAG ID plus run ID, which worked until 2 DAGs collided on a naming convention. Another built a little Postgres table called something like job_checkpoints that nobody documented and everybody was scared to touch.

The state store makes the fix boring, which is the highest compliment I give infrastructure. Submit the job, write the job ID to task state, and on retry check for an existing ID before resubmitting. Same idea for watermarks: persist the last processed offset so a retry resumes where it left off instead of re-reading the whole source.

Key takeaway: the state store gives Airflow a native answer to "what did this task already do?" That question sits under every duplicate-write incident you've ever debugged. Learn it before you learn the Java SDK.

Where I'd be careful

The default backend is the metadata database. That database already carries your scheduler, and it's the thing that falls over first when a deployment grows. Stuffing high-churn checkpoints into it for hundreds of tasks is a load decision, so make it on purpose. Set retention. Use clear_on_success where state only matters for in-flight recovery. Consider the worker-side backend if you're checkpointing aggressively.

Also, a state store doesn't make your pipeline idempotent. It gives you a place to record progress. Whether your writes are safe to replay is still your problem, and it's the same problem it was in 2015. If that concept feels shaky, go through idempotent pipeline design before you reach for the new API.

The Airflow Task SDK for Java and Go

This is the feature that'll get the conference talks. The Language Task SDK adds a Coordinator layer: JavaCoordinator routes tasks to JVM runtimes, and ExecutableCoordinator runs self-contained native binaries like Go.[2] Your DAG stays in Python. Individual tasks are declared with a @task.stub(queue=...) decorator, and the actual implementation lives in a JAR or a Go binary on standard Airflow workers.[1]

The Python side looks roughly like this; the body lives elsewhere, and the queue routes the task to workers with the right runtime.

from airflow.sdk import dag, task
@dag(schedule=None)
def orders_pipeline():
@task.stub(queue="java-workers")
def enrich_orders(): ...
@task
def publish_summary():
print("enrichment finished")
enrich_orders() >> publish_summary()
orders_pipeline()

Variables, Connections, XCom, and logging work in both SDKs. XCom values are stored as JSON in the metadata database, which acts as the serialization boundary between Python and the native runtimes.[1] The SDKs ship outside the apache-airflow pip package: Java as a Maven/Gradle dependency, Go as a module plus a coordinator binary.[1] Java workers need JRE 17 or later.[2]

The project's pitch: "Airflow is Python-first, but production business logic is often Java-first."[1] Fair. I've wrapped enough JARs in BashOperator calls to know the pain; you lose logging, you lose clean XCom, and the stack trace you get at 2am is whatever java -jar felt like printing to stderr.

The status label is doing a lot of work

The Coordinator layer and both SDKs are marked experimental in 3.3.0 and "may change in future versions based on user feedback."[2] The 3.3.2 notes include fixes for IPC short reads "crashing the subprocess or hanging the supervisor."[4] That's a normal bug for a young subsystem, and it's also a bug that would ruin your weekend in production.

My read: if you have mature Java business logic that you're currently shelling out to, pilot it on a non-critical DAG. If you're thinking about writing new task logic in Go because it sounds cool, don't. Every runtime you add is another build pipeline, another dependency tree, another thing on-call has to understand. Language sprawl costs engineer hours, and engineer hours are the most expensive line item you have.

Airflow Retry Policy Goes Pluggable

Before 3.3, retries were a count and a delay. The default retry_delay is 300 seconds, with max_task_retry_delay capping backoff at 24 hours.[5] Every failure got the same treatment. An expired credential got retried 3 times, 5 minutes apart, so you learned about it 15 minutes late with 3 identical stack traces.

3.3 adds ExceptionRetryPolicy built from RetryRule objects. Each rule names an exception type, an action (RETRY or FAIL), an optional custom delay, and a reason. Rules are evaluated in order, first match wins, and if nothing matches the task's normal retry count applies.[5] Every decision gets written to the task log as Retry policy decision action=<action> reason=<reason>.[5]

That log line is the underrated part. Half of debugging is reconstructing why the system did what it did. A logged reason string saves the 20 minutes you'd otherwise spend reading operator source code.

How I'd set up rules on day 1:

  • Fail immediately on authentication errors and schema mismatches. Retrying those wastes compute and delays the alert.
  • Retry rate limits and 503s with a longer delay than the default, since hammering a throttled API 5 minutes later rarely helps.
  • Leave everything else on the default count until your logs tell you otherwise.

You don't need anything fancier than that. Classifying your own failure modes is the actual skill here, and it's a skill most teams skip because blanket retries hide the problem.

Analysts Are Slowing the Store Down

> Analytics queries against the production database are slowing the live application. Move analytics onto its own warehouse fed from the database's change log, while a merchant dashboard shows new orders within fifteen minutes on a path of its own.

+ Source
+ Transform
+ Storage
+ Quality
+ Consumer
+ Queue
Bronze
Silver
Gold
Custom
Pipeline Architecture
Sketch the architecture.

Click or drag a node from the toolbar above. Right-click the canvas for the full menu.

Drag from a node's right port to another node's left port to wire data flow.

Airflow Asset Partitioning Gets Rollups and Fan-Out

3.3 adds 3 partition mappers that connect upstream asset events to downstream partitioned DAG runs:[2]

  • RollupMapper: many upstream partitions into one downstream run (daily partitions rolling into a weekly run).
  • FanOutMapper: one upstream partition into many downstream runs.
  • FixedKeyMapper with SegmentWindow: categorical rollups.

These compose with time windows (day, week, month, quarter, year) and wait policies, either WaitForAll or MinimumCount(n), which decide when a partitioned run fires.[2]

Fan-out has a safety valve. partition_mapper_max_downstream_keys defaults to 1,000 and can be set globally or per mapper. Exceed it and Airflow logs a fan-out exceeded event and creates no DAG runs for that asset event.[6]

AIRFLOW__SCHEDULER__PARTITION_MAPPER_MAX_DOWNSTREAM_KEYS=1000

Read that behavior twice. Blow past the cap and zero runs happen, with a log event as your only signal. That's a correct design choice (a runaway fan-out that creates 50,000 runs would be worse) but it's also a textbook silent-drop failure. If you use FanOutMapper, alert on that event. I once spent 6 months not noticing a job was dropping 40% of records. Don't be me.

Also worth knowing: 3.3 adds a PartitionedAtRuntime timetable for cases where the partition key only exists once the task runs, like a watermark pulled from a source system.[2] Pair that with the asset state store and you have a native watermark pattern that used to need a side table.

Partition design is data modeling wearing an orchestrator costume. Choosing between many small runs and fewer large ones comes down to grain, cardinality, and SLAs. Those decisions transfer to Dagster, to dbt, to whatever ships in 2029.

What Airflow 3.3 Means for Data Engineer Orchestration Skills in 2026

Here's my analysis, separate from the facts above. Interviewers lag releases by a year or more, and with Airflow 3 at 26% of users overall (48% among Astronomer's own customers), most loops in late 2026 will still assume 2.x habits.[7] Nobody's going to fail you for not knowing @task.stub syntax.

What they will probe is the stuff underneath these features, which has been interview material forever:

  • Failure recovery. "Your worker crashes mid-job. What happens on retry?" The state store is one answer; understanding why duplicates happen is the real answer.
  • Error classification. "Which failures should retry and which should page someone?" Retry policies just give that judgment a config surface.
  • Partition grain and cardinality. "How would you handle a daily source feeding a weekly aggregate?" Rollups and fan-out are the Airflow spelling of an old question.
  • Language boundaries. "When would you not add another language to the stack?" The honest answer involves on-call load and build complexity.

If you mention 3.3 features in an interview, tie them to the problem they solve. "I'd store the Spark job ID in task state so a retry checks before resubmitting" beats "Airflow 3.3 added AIP-103" every single time. The Airflow interview questions page covers the fundamentals these features build on, and the data engineering system design guide covers the pipeline architecture framing interviewers actually use.

Concepts transfer across tools; tool knowledge doesn't transfer across concepts. Airflow just shipped 4 features that are all old concepts with new APIs. The engineers who already understood idempotency, watermarks, and failure taxonomy will pick 3.3 up in an afternoon.

If you want reps on explaining these tradeoffs out loud, which is a different skill from knowing them, the mock interview simulator on DataDriven is the best practice I've found for that. It forces you to articulate the "why" under time pressure, which is where most people with real experience still get downleveled.

What to Watch Next

Based on what's shipped so far:

  • Upgrade to 3.3.2 if you're on 3.3.0. The 3.3.1 secret-handling fixes aren't optional for multi-team deployments.[3]
  • Adopt the state store first. It replaces workarounds you already maintain, and it's the one headline feature not labeled experimental. Watch metadata DB load and set retention from the start.
  • Write 3 to 5 retry rules for your noisiest DAGs. Fail fast on auth and schema errors; that's the cheapest reliability win in the release.
  • Alert on partition fan-out exceeded events before you ship a single FanOutMapper to production.
  • Hold on the Java/Go SDK unless you have existing JVM logic you're shelling out to today. Watch the next minor release for whether the experimental label comes off and whether the IPC layer stabilizes.
  • Watch the adoption numbers. 84% of users planning to upgrade is intent, not migration.[7] Next year's survey will tell you whether 3.x knowledge has become a baseline expectation in hiring.

The orchestrator keeps getting better at remembering what happened. Your job is still figuring out why it happened. For the rest of the prep, the data engineering interview prep guide is where I'd start.

References

  1. Apache Airflow, "Airflow 3.3.0: Stateful Tasks and Multi-Language Support," July 6, 2026. airflow.apache.org
  2. Apache Airflow, "Release Notes, Airflow 3.3.0." airflow.apache.org
  3. Apache Airflow, "Release Notes, Airflow 3.3.1," August 12, 2026. airflow.apache.org
  4. Apache Airflow, "Release Notes, Airflow 3.3.2," September 17, 2026. airflow.apache.org
  5. Apache Airflow, "Tasks," core concepts documentation. airflow.apache.org
  6. Apache Airflow, "Configuration Reference, Airflow 3.3.2." airflow.apache.org
  7. Astronomer, "State of Apache Airflow 2026 Report," survey of 5,800+ data practitioners, September 15 to November 20, 2025. astronomer.io

Apache Airflow 3.3Airflow Task SDK Java GoAirflow state storeAirflow asset partitioningAirflow retry policydata engineer orchestration 2026
02 / Why practice

Try the actual problems

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes