Data Engineer Roadmap (2026)

6 Stages in Order

The data engineer roadmap is 6 stages in order, from SQL and Python to interview practice. Dimensional modeling comes before depth in 1 warehouse on 1 cloud, and orchestration is the last stage before interviews. From scratch it takes 12 to 16 months at 10 to 15 focused hours a week, less with an analytics or software background, and every stage ends in something you build that proves it's done.

Last updated: Proudly published by: Jeff Wahl12 min read

The data engineer roadmap in order

The data engineer roadmap has 6 stages, taken in this order. SQL to fluency and Python for data work come first, then dimensional modeling, followed by 1 warehouse and 1 cloud in depth. Orchestration with its failure modes comes next, and interview practice closes the roadmap as a stage of its own. From scratch, at 10 to 15 focused hours a week, it takes 12 to 16 months; a data analyst or a software engineer can finish in a fraction of that.

The order matters more than the number of tools. Each stage assumes the one before it: you cannot state the grain of a fact table without reading SQL fluently, and you cannot make a pipeline safe to rerun without knowing where its tables land. Most published roadmaps are tool inventories of 30 or 40 logos. They give no order and no point at which you are done. Here every stage ends in a checkable output, something you build that either works or does not, and that output is the only proof the stage is finished.

Every day, a working data engineer writes SQL and Python against a warehouse while an orchestrator runs the jobs. Spark and Kafka show up on some teams and not others. Both are easier to learn once those 4 are solid, and the same goes for Flink or Terraform. Hiring loops follow the same shape: the coding rounds are SQL and Python, the design rounds ask for a data model or a pipeline, and Spark is added when the job description names it.

Prepare for the interview
01 / Open invite
02min.

Know Data engineer roadmap the way the interviewer who asks it knows it.

a Data engineer roadmap query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1SELECT user_id,
2 COUNT(*) AS sessions
3FROM events
4WHERE ts >= NOW() - INTERVAL '7 day'
5
Execute your solution0.4s avg.

How long the data engineer roadmap takes for your background

Plan on about 14 months from scratch, 12 to 16 depending on how many weeks work and life take back, at 10 to 15 focused hours a week. Focused means solving problems whose output gets checked, not watching videos. These are planning estimates built from stage lengths, not a measured average, and your pace will differ.

Background shortens the start of the roadmap far more than its end. A data analyst who writes SQL daily skips most of stage 1 and lands around 6 to 8 months. A working software engineer needs around 2 weeks of SQL and skips the Python stage, finishing in 3 to 4. A data scientist keeps Python but starts from zero at stage 3, for 4 to 6 months; a bootcamp graduate with broad but shallow coverage needs 3 to 5 months of depth on 1 warehouse and interview practice.

Months on the data engineer roadmap by background, each bar split into SQL and Python, then modeling onward: from scratch 14 months, 64% of it modeling onward; data analyst 7 months, 71%; software engineer 3.5 months, 86%Months on the data engineer roadmap by background, each bar split into SQL and Python, then modeling onward: from scratch 14 months, 64% of it modeling onward; data analyst 7 months, 71%; software engineer 3.5 months, 86%

So the faster your track, the larger the share that sits from modeling onward: 64% from scratch and 86% for a software engineer. An experienced engineer who budgets 2 weeks for the whole switch usually runs out of time at stage 3, not stage 1, so build the calendar around modeling and orchestration.

Stage 1: SQL, the first programming language to learn

SQL comes first because every later stage is written in it. Learn SELECT with joins and GROUP BY with HAVING, and once those are solid, move on to CTEs and window functions. After that, get NULL handling and conditional aggregation right, and know why COUNT(*) and COUNT(col) return different numbers. The time transfers to any data role: in the 2025 Stack Overflow Developer Survey, 61.3% of professional developers reported using SQL, more than the 54.8% who used Python.

You're done when you can write a window function query that handles ties correctly without looking anything up, explain why an INNER JOIN dropped rows you expected to keep, and deduplicate a table to the latest row per key. Ship 1 analysis on a real public dataset that ends in a query you can defend out loud.

Write SQL against a real database instead of reading about it, and let the queries that come back wrong show you where your picture of the data is off. The SQL practice problems check your output against the expected rows, and the SQL interview questions cover the patterns hiring loops return to. Ranking and running totals are there, and so are sessions and gaps and islands.

Where people stall: solve rate by stage on DataDriven practice problems

SQL
91%
Python
88%
Data modeling
51%
Pipeline design
34%

Share of members' attempts at each domain's problems that end in a passing solution, over all attempts

Stage 2: Python for data work

Python is the second language of the job, the glue around SQL. Learn the standard library for files (CSV, JSON, gzipped logs, nested records) and use dictionaries and generators for data too big to hold at once. In pandas, learn groupby and merge first; pivot can wait until those feel natural. Add enough error handling to tell a script from a job. A class with 3 methods and a test is about the ceiling of the object-oriented work the rounds ask for.

Skip algorithm grinding. Data engineering loops ask you to parse and reshape data and then check that it is valid, while puzzles about trees and graphs are the exception. You're done when you can ingest a messy file, reject the bad rows with a reason for each, and write the good ones to a warehouse table; ship exactly that script. The Python practice problems are built around parsing and transformation for this reason.

Stage 3: dimensional modeling, where most plans stall

Dimensional modeling is the stage most roadmaps treat as a footnote and the one that most often separates mid-level from senior candidates. It's the discipline of shaping tables for analysis: facts that record events, dimensions that describe them, and a grain that says exactly what 1 row means.

Learn the Kimball Group's 4-step dimensional design process: select the business process, declare the grain, identify the dimensions, then identify the facts. The order matters because a candidate who draws tables before stating the grain ends up with a fact table that mixes 1 row per order with 1 row per order line, and every sum on it is wrong.

Then learn slowly changing dimensions. Type 1 overwrites an attribute; Type 2 adds a new row with validity dates so history survives; the choice is a product question about whether past facts should report the old value or the new one. The Data Warehouse Toolkit by Ralph Kimball and Margy Ross (3rd edition, 2013) is the reference for the patterns. Ship a 5 table model for a business you understand, with the grain written above every fact table.

On the practice problems at datadriven.io, this is where people stall: members pass 91% of their attempts at SQL problems and 51% of their attempts at data modeling problems. Practise it aloud with the data modeling interview questions, and read how the data modeling interview round is run before your first loop.

A food delivery model that keeps courier history

fact_delivery
delivery_idPKBIGINT
courier_keyFKBIGINT
customer_idFKBIGINT
city_idFKINT
delivered_atTIMESTAMP
delivery_feeDECIMAL
dim_courier
courier_keyPKBIGINT
courier_idBIGINT
home_cityVARCHAR
valid_fromDATE
valid_toDATE
is_currentBOOLEAN
dim_customer
customer_idPKBIGINT
signup_dateDATE
segmentVARCHAR
dim_city
city_idPKINT
city_nameVARCHAR
countryVARCHAR

The grain is 1 row per completed delivery. dim_courier is a Type 2 dimension: a courier who moves city gets a new row with a new courier_key, so past deliveries keep the home city that was true when they happened, and joining on courier_id instead would rewrite that history.

Stage 4: which warehouse and cloud platform to learn first

Pick 1 warehouse and learn it past the surface. Snowflake and BigQuery both qualify, and so does Postgres. Learn how it stores and prunes data through partitioning and clustering, how to read a query plan, what a materialized view costs to keep fresh, and which dialect quirks change what is cheap. Surface familiarity with 3 warehouses rarely survives a follow-up question; depth on 1 lets you contribute in your first week.

Cloud works the same way. Choose 1 platform by the companies you target: AWS if you've no preference, Google Cloud where BigQuery is the warehouse, Azure where the employers run Microsoft Fabric. Stop at the depth an interview reaches: how data lands in object storage, how it's partitioned, what a query costs, and how a job authenticates to read a bucket. Everything past that you'll learn on the job.

Stages 3 and 4 overlap on purpose. Build your model in the warehouse you chose, so the grain you declared runs into real partitions and query plans, and a real bill.

Stage 5: orchestration and pipelines that are safe to rerun

Orchestration turns a script into a pipeline that runs on a schedule and waits for the tasks it depends on. When a task fails, the pipeline retries it and sends an alert, and it can backfill past dates. Learn Apache Airflow first, since it is the orchestrator most teams run and most interviewers assume, and its workflows are plain Python. Dagster and Prefect are real alternatives and quick to pick up once Airflow's model makes sense.

The idea to master is idempotency: a task run twice for the same date must leave the same result as a task run once. Retries and backfills rerun tasks as a matter of routine, so a load that's not idempotent corrupts data during normal operation. The Airflow best practices guide says to treat a task like a transaction in a database and warns that an INSERT during a rerun can duplicate rows.

The same day of orders loaded 3 times: with INSERT the table grows from 1,000 to 2,000 to 3,000 rows, 2,000 of them duplicates, while with INSERT OVERWRITE it holds 1,000 rows after every runThe same day of orders loaded 3 times: with INSERT the table grows from 1,000 to 2,000 to 3,000 rows, 2,000 of them duplicates, while with INSERT OVERWRITE it holds 1,000 rows after every run

Take a daily load of 1,000 orders. Run it 3 times for the same day with a plain INSERT and the table holds 3,000 rows for that day, 2,000 of them duplicates, so every revenue sum is tripled. Write it as a delete and reload, or an INSERT OVERWRITE of that date's partition, and it holds 1,000 rows after every run.

The stage also covers late data and schema drift. It explains why exactly-once delivery usually means at-least-once plus deduplication, and the data pipeline interview questions draw on all of it. Ship a DAG that loads something on a schedule, survives a deliberate failure and backfills a week without a duplicate row.

Stage 6: interview practice as its own stage

Skill doesn't turn into offers on its own. Interviews reward recall under a clock and the ability to narrate a tradeoff out loud, and both are separate from doing the work. Give them their own weeks: timed problems with the output checked, modeling questions answered aloud, and whole loops run end to end.

A data engineering loop usually has a SQL round and a Python round, then a data modeling round and a pipeline or system design round, plus a behavioral round. The mix shifts by level. The data engineer interview prep guide breaks down what each round covers, and a data engineer mock interview is the closest rehearsal. You are done when a full timed loop no longer surprises you.

What to skip, and when a portfolio project helps

Skip the tool inventory. Spreading the months before your first application across Kafka, Spark, Flink, dbt, Terraform and 3 clouds leaves each one thin, and thin knowledge fails follow-up questions. Learn Spark when a target job names it; its DataFrame API reads like the SQL and pandas you already know.

Skip streaming before batch. Teams usually run batch pipelines first and add streaming where the latency pays for it, and every streaming idea (late events, watermarks, exactly-once) is easier once batch failures are familiar. Skip certifications as a hiring strategy too, since the rounds probe problem solving rather than vendor knowledge.

A portfolio project helps in the last third of the roadmap, not the first month. Built on shaky SQL, it is a liability in the interview where someone asks you to explain it. Build it after stage 5, on a real dataset, and write up the design choices: the grain, how it reruns, what it costs. The writeup gets read more closely than the code.

Data engineer vs analyst, analytics engineer, data scientist and ML engineer

RoleWhat it ownsCore skillsHow this roadmap changes
Data engineerPipelines, models, freshness and correctness of the dataSQL, Python, modeling, orchestrationAll 6 stages, in order
Analytics engineerTransformations and metric definitions inside the warehouseSQL, dbt, modelingStages 1, 3 and 4 in depth, with lighter orchestration
Data analystAnswers, dashboards and the questions behind themSQL, business context, communicationStage 1 and the modeling basics, then statistics and visualisation
Data scientistModels, experiments and their evaluationStatistics, Python, machine learningStages 1 and 2, then statistics in place of stages 3 to 5
ML engineerTraining and serving infrastructure for modelsPython, distributed systems, ML toolingStages 2 and 5 in depth, plus model serving

Is data engineering still worth entering in 2026?

AI assistants now draft routine SQL and boilerplate transformations well, so less of a junior role goes to hand-writing a straightforward pipeline than it once did. What they do not remove is the judgment around the code: which grain a table has, whether a load is safe to rerun, what a query will cost, who gets paged when a table is late. Interviews have moved the same way and ask why more often than what.

The US Bureau of Labor Statistics has no separate data engineer occupation. Its closest, database architects, earned a median $139,500 in May 2025, and it projects 4% growth for database administrators and architects together from 2025 to 2035. Stages 3 to 5 of the roadmap, the judgment layer, are the part that gains value as assistants take over the typing.

How to follow the roadmap while working a full-time job

Beside a full-time job, the constraint is attention more than hours. The usual failure is spending month 4 rereading stage 1 because it feels productive. Give every week a checkable output: a query that returns the right rows, a model with its grain written down, a DAG that survives a deliberate failure.

Keep a data-driven log with 3 columns. Record the problem and whether it passed, and in the third column write what you got wrong. Rework the misses a week later without notes. Once window functions stop being effortful, start the data modeling practice problems even if SQL still feels unfinished; modeling practice keeps your SQL sharp, and the reverse isn't true. At 10 hours a week, 45 focused minutes on weekdays plus a longer weekend session beats 1 lost Saturday.

Data engineering certifications: when one is worth it

A certification rarely decides a data engineering offer, because the loop probes problem solving rather than vendor knowledge. It helps in 2 cases: a role or a consultancy that names it as a requirement, and a career switcher who needs a resume line that shows cloud exposure.

If you take one, take a current exam. Microsoft retired the Azure Data Engineer Associate certification (DP-203) on March 31, 2025. Its current data engineer certification is Fabric Data Engineer Associate (DP-700), which covers SQL and PySpark as well as KQL. The AWS Certified Data Engineer Associate (DEA-C01) is 65 questions in 130 minutes, and AWS recommends 2 to 3 years of data engineering experience before it, which places it after this roadmap rather than inside it.

Data engineer roadmap FAQ

How long does it take to become a data engineer?+
Plan on 12 to 16 months from scratch at 10 to 15 focused hours a week. A data analyst who writes SQL daily needs about 6 to 8 months, a working software engineer 3 to 4, a data scientist 4 to 6 and a bootcamp graduate 3 to 5. These are planning estimates, and the hours count only when the work is checked: solved problems, working pipelines, models with their grain written down.
What is the data engineer roadmap in order?+
Start with SQL to fluency and Python for data work. Stage 3 is dimensional modeling, and stage 4 is depth in 1 warehouse on 1 cloud. Orchestration follows with its failure modes, and interview practice closes the roadmap as a stage of its own. Each stage assumes the one before it, and the most common sequencing mistake is starting on a tool from stage 4 or 5 before SQL is automatic.
Can I become a data engineer without a CS degree?+
Yes. None of the interview rounds asks for a degree: they check whether you can write SQL and Python and whether you can design a data model or a pipeline. A degree helps most at the resume screen; after that, what you can do in the rounds decides the outcome, so a portfolio project with a clear design writeup matters more for a candidate without one.
What's the fastest path if I'm already a software engineer?+
About 3 to 4 months. After 2 weeks making SQL fluent, you start at dimensional modeling and work through the rest of the stages in order. Your algorithm preparation is more than these loops need; the gaps to close are the grain of a table and the failure modes of scheduled jobs that rerun, which differ from those of a service.
Should I get a certification?+
Only when a target role names one or when you need proof of cloud exposure on a career switch. Rounds probe problem solving rather than vendor knowledge. If you take one, take a current exam: Microsoft retired DP-203 on March 31, 2025. Its current data engineer exam is DP-700 for Microsoft Fabric.
What does a data engineer earn in 2026?+
The US Bureau of Labor Statistics has no separate data engineer occupation. Its closest, database architects, earned a median $139,500 in May 2025, with the top 10% above $204,000. Pay at large tech companies runs higher, and the spread inside every level is decided in the interview loop.
Is data engineering being automated away by AI?+
No. Assistants draft routine SQL and boilerplate transformations well, which moves the valuable part of the job, and of the interview, to judgment: the grain of a table, whether a load is safe to rerun, what a query costs and who is paged when data is late. Stages 3 to 5 of the roadmap are that layer.
Is math required for data engineering?+
Very little beyond arithmetic and the set logic behind joins. Calculus and linear algebra belong to data science, as does inferential statistics. The hard parts of data engineering are precision about data and reasoning about systems that fail.
02 / Why practice

The candidate who gets the offer

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    5 problem shapes cover 80% of data engineer loops

    Dedup, sessionization, top-N-per-group, slowly-changing dimensions, partition tricks. Writing the shapes by hand turns the unfamiliar into pattern recognition

Related guides