Data Engineer Interview Questions by Round

Data engineer interview questions cover 5 rounds: the behavioral round and 4 technical ones on SQL, Python, data modeling and pipeline design. Companies that run Spark add a Spark round. Every question on this page was asked in a real data engineering interview and is answered the way interviewers want to hear it, with the traps that cost points in each round.

Last updated: Proudly published by: Jeff Wahl20 min read

What data engineer interviews ask

A data engineering loop asks questions from 5 rounds: 4 technical ones on SQL, Python, data modeling and pipeline design, plus behavioral, with a Spark round at companies that run Spark in production. Every technical round tests the same thing from a different side: whether the data you produce is correct when the input carries duplicates and NULLs, ties and skew, or late rows and reruns.

The practice catalogue on DataDriven holds {COUNT:total} questions asked in real data engineering interviews, tagged with {COUNT:companies} employers: {COUNT:sql} in SQL, {COUNT:python} in Python, {COUNT:data_modeling} in data modeling and {COUNT:pipeline_architecture} in pipeline architecture. Each round has its own full guide: SQL interview questions, Python interview questions, data modeling interview questions, data pipeline interview questions and PySpark interview questions.

What loses points is rarely syntax. In SQL it is a missed NULL, a join that changes the grain, or a ranking that breaks on ties. In Python it is a bare except, a whole file read into memory, or a nested loop where a dictionary lookup would do. In modeling it is drawing tables before saying what 1 row means. In design it's the right boxes with no plan for reruns and late events, or for a schema change. In Spark it is a stage stuck on 1 task that the candidate can't explain.

Prepare for the interview
01 / Open invite
02min.

Know Data Engineer Interview Questions the way the interviewer who asks it knows it.

a Data Engineer Interview Questions query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1SELECT user_id,
2 COUNT(*) AS sessions
3FROM events
4WHERE ts >= NOW() - INTERVAL '7 day'
5
Execute your solution0.4s avg.

How the data engineer interview loop is structured

Most loops open with a recruiter call and a technical screen. The screen is usually SQL in a shared editor, sometimes with a short Python task, and it decides whether you reach the full loop. The full loop then runs 4 or 5 interviews of about an hour each, sometimes on 1 day and sometimes across a week. It covers SQL in more depth along with Python, data modeling and pipeline or system design, plus a behavioral round. Some companies replace the live coding with a take-home assignment and a review of it.

Each round has a different interviewer listening for a different thing. The coding rounds check correctness on the cases the prompt does not mention. In modeling, the interviewer wants tables that answer the business questions at the grain you claimed, and in design, a pipeline that survives failure. The behavioral interviewer wants to hear whether you've done this work before and what you learned from it. A strong round rarely rescues a weak one, so preparation spreads across all of them; the data engineer interview prep guide lays out the order the rounds run in and how to plan for each.

SQL interview questions for data engineers

SQL

Find the second highest salary in each department.

Rank salaries within each department with DENSE_RANK() OVER (PARTITION BY department_id ORDER BY salary DESC) in a CTE, then keep the rows where the rank is 2. DENSE_RANK gives tied salaries the same rank and leaves no gap, so the answer is the second highest distinct salary and every employee who earns it. ROW_NUMBER would return 1 of 2 tied people at random, and RANK returns nothing for 2 when 2 people tie for first. Before you write the query, say whether a department with 1 employee should return no row or a NULL.

SQL

Keep only the latest row for each customer from a table of updates.

Number the rows within each customer_id with ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY updated_at DESC, update_id DESC) and keep row 1. The second ORDER BY column is the tie-breaker: without it, 2 updates in the same second come back in any order and the result changes between runs. GROUP BY with MAX(updated_at) finds the latest time but not the other columns of that row, and joining back to it returns both rows on a tie.

SQL

List the customers who have never placed an order.

Use NOT EXISTS with a correlated subquery, or a LEFT JOIN from customers to orders that keeps the rows where the order's key IS NULL. Avoid NOT IN on a column that can hold NULL: if the subquery returns a single NULL, NOT IN returns no rows at all. Interviewers ask this question to see whether you know that.

SQL

Find each user's longest streak of consecutive login days.

Deduplicate to 1 row per user and day first, then subtract ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY login_date) days from each date. Consecutive days share the same result, so group by user and that anchor date, count the days in each group and take the maximum per user. Skip the deduplication and a user who logs in twice on 1 day gets a streak that jumps a gap.

SQL

Compute a 7-day rolling average of daily revenue.

Aggregate to 1 row per day, then AVG(revenue) OVER (ORDER BY day ROWS BETWEEN 6 PRECEDING AND CURRENT ROW). ROWS counts rows, not days, so a day with no sales shrinks the window to fewer calendar days. Join to a calendar table that has every date first, filling missing revenue with 0, or use a RANGE frame of 6 days where the engine supports it.

SQL

A LEFT JOIN returns fewer rows than the left table has. Why?

A condition on the right table sits in the WHERE clause. Rows with no match carry NULL in every right-hand column, the WHERE condition isn't true for NULL, and those rows are dropped, which turns the LEFT JOIN into an INNER JOIN. Move the condition into the ON clause so it filters what joins rather than what survives.

SQL

Revenue by month doubled after you joined orders to another table. What happened?

The join fanned out: the other table has more than 1 row per order, such as 1 row per shipment or per promotion, so each order's amount is repeated once per match. Check the grain of both sides before joining, aggregate the many side to 1 row per order first, or sum from the orders table alone.

SQL

A daily query that ran in 2 minutes now takes 40. How do you find out why?

Read the query plan before changing anything. Look for a full scan where a partition filter should prune, or a function wrapped around the partition or clustering column that stops pruning. A join that multiplies rows before a filter is another cause, and so is a sort or shuffle that spills to disk. Then compare the input sizes with last month's, because a query that slows as its tables grow needs a different fix from one whose plan changed.

The SQL mistakes interviewers wait for

SQL interviewers choose data with the edge cases already in it, then wait to see whether your query handles them. The most common is a NULL in a list. The PostgreSQL documentation on subquery expressions states the rule: when no value in the list equals the left-hand value and at least 1 list value is NULL, NOT IN yields NULL, not true. A WHERE clause keeps only rows that are true, so the query returns nothing and raises no error.

The NOT IN demonstration runs on 3 customers and 2 orders, and 1 of the orders has no customer_id. NOT EXISTS finds the 2 customers who never ordered. NOT IN finds 0, because the NULL in the subquery makes every comparison unknown. Add WHERE customer_id IS NOT NULL to the subquery and the 2 methods agree again. The interview version of the question, users who never made a purchase, opens in the practice problem card, where your answer runs against hidden data.

The other traps follow the same pattern. COUNT(column) skips NULLs where COUNT(*) does not, so the 2 disagree on exactly the rows an interviewer seeded. Integer division truncates in PostgreSQL and SQL Server, so a ratio of 2 integer counts comes back as 0 unless 1 side is cast to a decimal. A filter on a timestamp column written as = '2026-01-31' matches only midnight. Say each assumption aloud as you write, then check it against the sample rows before you call the query done.

For more questions of this kind, all with answers, the SQL interview questions guide works through the round in full and the SQL practice problems run every query against a real database.

Run it: NOT IN vs NOT EXISTS when a NULL is in the list

SELECT 'NOT IN' AS method, COUNT(*) AS customers_without_orders
FROM customers
WHERE customer_id NOT IN (SELECT customer_id FROM orders)
UNION ALL
SELECT 'NOT EXISTS', COUNT(*)
FROM customers c
WHERE NOT EXISTS (
  SELECT 1 FROM orders o WHERE o.customer_id = c.customer_id
);

customers holds 3 rows and orders holds 2, 1 of them with a NULL customer_id. NOT IN returns 0 and NOT EXISTS returns 2. Add WHERE customer_id IS NOT NULL inside the NOT IN subquery and run it again.

ROW_NUMBER vs RANK vs DENSE_RANK: how ties change the answer

Ranking questions are where most window function answers go wrong, because the 3 ranking functions agree until 2 rows tie. Take 4 salaries in 1 department: 180, 150, 150, 120. ROW_NUMBER numbers them 1, 2, 3, 4 and breaks the tie by whatever order the engine reads the rows in. RANK gives 1, 2, 2, 4: the tied rows share 2 and the next rank skips to 4. DENSE_RANK gives 1, 2, 2, 3: the tied rows share 2 and the next rank is 3.

So the third highest salary depends on the function. A filter of = 3 returns 1 of the 150 rows with ROW_NUMBER, no row with RANK, and 120 with DENSE_RANK. A top 3 filter of <= 3 keeps 3 rows under the first function and 3 under the second, while the third keeps 4. Before writing the window, ask the interviewer whether tied people both count and whether "third highest" means the third distinct value.

When you need exactly 1 row per group, as in deduplication, ROW_NUMBER with a tie-breaking column keeps the choice stable. DENSE_RANK returns the Nth distinct value, and RANK fits a competition ranking where a tie for second means nobody is third.

Try it on a question from real interviews: the second highest latency for each API method, where several calls can tie.

Second Highest Latency by Method

> Find the second-highest latency API endpoint in each HTTP method group. If multiple endpoints share the highest latency, the second-highest is the next unique latency value below that. Show method, endpoint, and latency.

Ranking functions on salaries 180, 150, 150, 120

ROW_NUMBER
  • Ranks given1, 2, 3, 4
  • Tied rows getDifferent numbers
  • After a tieNext number
  • Rows where rank = 3150 (1 row)
  • Rows kept by rank <= 33
  • Best for1 row per group
RANK
  • Ranks given1, 2, 2, 4
  • Tied rows getThe same number
  • After a tieSkips ahead
  • Rows where rank = 3No rows
  • Rows kept by rank <= 33
  • Best forCompetition ranking
DENSE_RANK
  • Ranks given1, 2, 2, 3
  • Tied rows getThe same number
  • After a tieNext number
  • Rows where rank = 3120 (1 row)
  • Rows kept by rank <= 34
  • Best forNth distinct value

What the Python round asks data engineers

The Python round for data engineers is about moving and reshaping records, not puzzles. Expect to parse a file or a JSON payload, deduplicate, group, sessionize events by a time gap, or pull every page of an API without losing or repeating a record. Most answers need nothing beyond the standard library, mainly dict, set, collections.Counter, heapq and itertools.groupby, plus generators.

Interviewers listen for 3 habits. A 50 GB file is read 1 line at a time, never into a list, and you can say how much state your solution keeps. A bad record is caught with a specific exception and set aside with its reason, and the job keeps a count, so 1 malformed line neither crashes the job nor disappears silently. A lookup inside a loop goes through a dictionary or a set, which turns an O(n²) answer into O(n).

Many companies allow pandas, and some expect it, but be ready to write the same logic without it, because the follow-up is often "now the data does not fit in memory". The Python interview questions guide has the full set of questions for this round.

Python, data modeling, pipeline and Spark interview questions

Python

Return the 10 most active users from a 50 GB log file.

Stream the file line by line, parse the user id from each line, and count with a Counter, so memory grows with the number of distinct users rather than the file size. Counter.most_common(10) or heapq.nlargest returns the top 10 without sorting every user. If even the distinct users do not fit in memory, hash each user id into 1 of N temporary files, count each file separately and merge the per-file top 10s.

Python

Load every record from a paginated API that rate-limits you.

Follow the cursor or next-page token the API returns. Page numbers you compute yourself skip or repeat records when data changes during the load. On a 429 or a 5xx response, retry with exponential backoff and honor the Retry-After header. Write each page keyed by record id so a retried page upserts instead of duplicating, and save the cursor after each page so a crash resumes where it stopped.

Python

Group click events into sessions that end after 30 minutes of inactivity.

Sort the events by user and timestamp, walk them in order, and start a new session whenever the user changes or the gap since that user's previous event exceeds 30 minutes. Give each session an id from the user id and its first timestamp. Ask whether the input is already sorted and whether events can arrive late, because both change the solution.

Data modeling

Design the warehouse model for a ride-sharing company.

Start with the business process and the grain: a trips fact at 1 row per completed trip, measuring fare and distance along with duration. Add dimensions for rider and driver, plus vehicle and location, and a date dimension. Keep driver attributes such as rating tier as SCD Type 2 so past trips report the tier the driver had at the time. Payments and cancellations are different processes at different grains, so they become their own fact tables sharing the same dimensions.

Data modeling

When do you use SCD Type 1, Type 2 and Type 3?

Type 1 overwrites the attribute. Use it for corrections and for anything nobody reports history on. Type 2 closes the current row and inserts a new version with valid_from and valid_to, for attributes whose history changes past results, such as a customer's region. Type 3 adds a column for the previous value, for the rare case where reports need only the prior value side by side. Decide per attribute, not per table.

Pipeline design

How do you make a daily pipeline safe to rerun?

Make every run write a fixed slice of data and replace it: each run reads the partition for its own logical date and overwrites the matching output partition, or merges on a business key. Never append with a plain INSERT and never compute the date from the current time inside the task, or a retry after a partial failure leaves duplicates and a rerun next week processes the wrong day.

Pipeline design

Events arrive up to 2 days late. How does your daily aggregate stay correct?

Partition the raw data by event time, and on each run recompute not only yesterday but a lookback window covering the last 3 days, overwriting those partitions. Track how late events actually arrive so the window can be sized from data, and decide with the consumers what happens to events later than the window: a periodic backfill or a documented cutoff.

Spark

1 task in a Spark join runs for an hour while the rest finish in a minute. Why?

Data skew. A single join key holds a large share of the rows, and the shuffle sends all of them to the same task. The hot key is often a null or default value, or 1 very large customer. Check the stage's task durations and shuffle read sizes in the Spark UI to confirm it, then let adaptive query execution split the skewed partition, broadcast the small side if it fits, filter out null keys, or salt the hot key by hand.

Spark

What is the difference between repartition and coalesce?

repartition does a full shuffle and can raise or lower the partition count, giving evenly sized partitions, optionally by a column. coalesce only merges existing partitions without a shuffle, so it's cheaper but can only lower the count and can leave partitions uneven. Use coalesce to cut the number of output files at the end of a job and repartition when the next step needs even work.

How to answer data modeling interview questions

The data modeling round names a business and asks for its tables. It might be a marketplace or a rideshare app, a payments ledger or a streaming service. Most good answers follow the 4 decisions of Kimball's dimensional design process in order: select the business process, declare the grain, identify the dimensions, identify the facts. The Kimball Group defines the grain as exactly what a single fact table row represents, and recommends the atomic grain, the lowest level the process captures, because it answers questions nobody has asked yet.

Say the grain out loud before drawing anything: "1 row per completed trip" or "1 row per order line". Every later choice follows from that sentence, and most modeling mistakes are a table whose rows mean 2 things. Then name the questions the model must answer and check each against your tables as you go.

Expect the interviewer to press on each choice you made. They will ask which attributes need SCD Type 2 and how a fact row joins to the version that was true at the time, why you chose a star schema with conformed dimensions over a normalized design or 1 wide table, and what happens to the model when the business adds a new product line or a second currency. The data modeling interview questions guide covers each of these with full answers.

Users Without Purchases

> The growth team is sizing the conversion gap: how many registered users have never completed a purchase?

SCD Type 2: why the validity interval is half-open

SCD Type 2 keeps history by closing the current row of a dimension and inserting a new version, each with a valid_from and a valid_to. A fact row then joins to the version that was current when the event happened. Interviewers ask for that join, and the boundary is where it breaks.

Customer 42 lived in Austin from day 1 and moved to Denver on day 10, so version 1 ends on day 10 and version 2 starts on the same day. An order for $250 placed on day 10, joined with order_ts BETWEEN valid_from AND valid_to, matches 2 versions, because BETWEEN includes both ends. The order is counted twice and its revenue reads $500.

Customer 42's 2 SCD Type 2 versions, Austin from day 1 to day 10 and Denver from day 10, with an order for $250 on day 10: a BETWEEN join includes both ends and matches 2 rows, a half-open join excludes the end and matches 1 rowCustomer 42's 2 SCD Type 2 versions, Austin from day 1 to day 10 and Denver from day 10, with an order for $250 on day 10: a BETWEEN join includes both ends and matches 2 rows, a half-open join excludes the end and matches 1 row

Join on a half-open interval instead: order_ts >= valid_from AND order_ts < valid_to. Every moment then belongs to exactly 1 version, and the order counts once. Give the current row a far-future valid_to such as 9999-12-31 rather than NULL, so the same condition covers it without a COALESCE. More questions on history tables are in the SCD interview questions.

How to answer a data pipeline design question

A typical prompt: "Replicate the orders table from our production PostgreSQL database into the warehouse, with dashboards no more than 15 minutes behind." Start with the numbers the design depends on, rows changed per day and how fresh the data must be, then find out whether deletes matter and who reads the result. Ask for the ones the prompt leaves out.

Then walk the data from source to consumer. Capture changes from the database's write-ahead log with change data capture rather than polling an updated_at column, which misses hard deletes and any update that forgets to set it. Publish the changes to a Kafka topic keyed by order_id, so every change to 1 order lands in 1 partition and stays in order. A Spark job merges each batch into a lake table on order_id, applying a change only when its log position is newer than the stored one. Under that rule a replay or a duplicate leaves the table as it was. Records that fail to parse go to a dead letter queue with the reason attached, and a quality gate runs before the table reaches the dashboard.

The design is the easy half. The interviewer then asks what happens when things fail: the merge job dies halfway, the source adds a column, an event arrives 2 days late, the topic must be replayed from last week, or the table needs a backfill from a fresh snapshot. Answer each with the mechanism that handles it, and name the trade-off you accepted. The data pipeline interview questions guide and the system design round guide cover more prompts of this kind.

A CDC pipeline from PostgreSQL to a dashboard

Capture
Process
Serve
PostgreSQL
orders db
CDC
orders cdc
Kafka
orders.changes
Spark
merge orders
SLA< 15minIDEMPOTENCYMERGE on order id, newer log position wins
dead letter queue
orders dlq
Iceberg
lake.orders
dbt tests
orders checks
Looker
revenue dashboard

Changes are read from the PostgreSQL write-ahead log and keyed by order_id in Kafka. They're merged into an Iceberg table where a newer log position wins, so replaying a change or receiving it twice leaves the table unchanged. Rows that fail to parse go to a dead letter queue, and dbt tests gate the table before the dashboard reads it, 15 minutes behind the source at most.

Idempotent pipelines: what happens when a load runs twice

Every orchestrator retries failed tasks, and every team reruns a day after fixing a bug, so interviewers ask what your pipeline does when the same load runs twice. The answer they want is "the same as when it runs once". A pipeline with that property is idempotent.

Take a daily load that writes 50,000 rows into the day's partition. If it appends with a plain INSERT and the scheduler retries it after a timeout, the partition holds 100,000 rows, 50,000 of them duplicates, and every sum downstream is double. If it overwrites the day's partition, or merges on a business key, the partition holds 50,000 rows after the first run, and neither the retry nor any later rerun changes that count.

Rows in 1 day's partition after the first run and after a retry: with INSERT the retry leaves 100,000 rows, 50,000 of them duplicates shown in red; with an overwrite of the partition both runs leave 50,000 rowsRows in 1 day's partition after the first run and after a retry: with INSERT the retry leaves 100,000 rows, 50,000 of them duplicates shown in red; with an overwrite of the partition both runs leave 50,000 rows

The Apache Airflow best practices give the same rules. Replace INSERT with an upsert on rerun. Read and write a specific partition named by the run's data_interval_start rather than the latest data. Never call now() inside a task, because each run would then compute something different. Say those 3 rules in a design round and the follow-up questions about retries and backfills answer themselves.

Spark interview questions: shuffles, joins and skew

A Spark round checks whether you know where a job spends its time, and the answer is nearly always the shuffle. Narrow transformations like filter or select work inside each partition, and so does withColumn. Wide ones such as groupBy or join move rows across the cluster and cut the job into stages, and distinct does the same. Skew and disk spills show up in those stages, and that's where the slow tasks are.

Know the defaults, because the questions lean on them. Spark broadcasts the smaller side of a join automatically when it's under 10 MB, the default of spark.sql.autoBroadcastJoinThreshold, and a shuffle produces 200 partitions unless spark.sql.shuffle.partitions says otherwise. Adaptive query execution, on by default since Spark 3.2.0, coalesces small shuffle partitions and splits skewed ones while the job runs.

The skew rule has 2 conditions, and interviewers like to test the second. The Spark performance tuning guide treats a partition as skewed only when it's larger than 5 times the median partition and also larger than 256 MB. With a median of 40 MB, 5 times the median is 200 MB, but a partition still has to pass 256 MB before it's split. Below that, a slow task is either acceptable or a case for salting the hot key by hand. The PySpark interview questions guide goes deeper on joins and skew and on reading the Spark UI.

Behavioral questions for data engineers

The behavioral round asks about work you have done, and for data engineers the questions cluster around a few themes: a pipeline that broke and how you found out, a time a number on a dashboard was wrong, a disagreement with a stakeholder about how a metric is defined, a project with unclear requirements, and a technical trade-off you made and would make differently now.

Describe the situation and what you did, then give the result in the units of the work. Count the hours data was late and the tables or consumers affected, or say how much a change saved in run time or cost. Then say what you changed so it couldn't happen again, such as a freshness check, a contract with the upstream team or a rerun procedure.

Prepare 5 or 6 stories that each cover several themes, and write down the numbers in each before the interview. The behavioral round guide has the full list of questions and how to structure each answer.

How the questions change from entry-level to senior

The topics stay much the same across levels; what the interviewer expects from your answer changes. At entry level, the bar is correct SQL and Python on the edge cases, a clean star schema and a clear account of a project you built. The design round, if there's one, asks for a simple batch pipeline. The entry-level data engineer interview questions focus on that loop.

At mid level, you're expected to own a pipeline end to end. That means choosing the grain and making the job idempotent, then testing the output and explaining how you would know it broke. At senior level, the design round carries the most weight. You're expected to ask for the numbers before drawing, name the failure modes of each component without being prompted, weigh cost against freshness, and adapt when the interviewer changes a requirement halfway through. The senior data engineer interview questions are pitched at that bar.

How to practice data engineer interview questions

Practice under the conditions of the round. Give yourself about 20 minutes for a SQL or Python question and 45 for a modeling or design question, talk through your plan before you write, and check your answer against the edge cases before you run it. The habit the interview rewards is finding your own bug before the interviewer points at it.

Spend the most time on SQL, because it appears in nearly every loop and usually in the screen that decides whether you reach the rest. Then work through the rounds your target companies run, using the data engineer interview prep guide to plan the order. In the final week, run full rounds with a timer in a data engineer mock interview, where an answer that sounds right but would fail on the data gets the same follow-up a real interviewer would ask.

Data engineer interview questions FAQ

How many rounds does a data engineer interview have?+
Most loops start with a recruiter call and a technical screen. After those come 4 or 5 rounds covering SQL, Python, data modeling and pipeline or system design, plus a behavioral round. Companies that run Spark in production often swap in or add a Spark round, and some replace the live coding with a take-home assignment.
What SQL questions are asked in data engineer interviews?+
Expect top N per group or the second or Nth highest value, keeping the latest row per key, customers with no orders and consecutive-day streaks. Running totals and rolling averages come up as well, and so does diagnosing a slow query. Each one tests window functions and joins, and how you handle NULLs, on data with duplicates and ties.
Do data engineer interviews include LeetCode-style algorithm questions?+
Some companies run a general coding round with data structures and algorithms, but most data engineering Python rounds manipulate data instead: parsing a large file and deduplicating records, or sessionizing events and pulling a paginated API. Nearly all of them come down to dictionaries and sets, with sorting, heaps or a generator where the data calls for one.
Do all data engineer interviews include a Spark round?+
No. Companies whose pipelines run on Spark or Databricks often include a Spark round on DataFrames and joins, with questions on shuffles and skew. Companies built on Snowflake or BigQuery usually test the same ideas through SQL and the design round instead.
How do I answer a data modeling question that has no single correct answer?+
State the grain in 1 sentence before you draw, list the questions the model must answer, then choose the facts and dimensions and decide how much history each dimension keeps. Interviewers accept several designs. They reject a design whose author cannot say what 1 row means or defend a choice when pressed.
What do interviewers look for in the pipeline design round?+
A design that meets the stated freshness and volume and has an answer for every failure: for reruns and duplicates, for late or out-of-order events, for a schema change upstream, and for a backfill. Naming the trade-off behind each choice matters more than naming tools.
How is a data engineer interview different from a data analyst interview?+
Both test SQL, but a data engineer loop asks you to build and run the data: Python for ingestion and transformation, modeling for the warehouse, and a design round about pipelines that stay reliable at scale. An analyst loop spends that time on metrics and experiments, and on business cases.
02 / Why practice

The candidate who gets the offer

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    5 problem shapes cover 80% of data engineer loops

    Dedup, sessionization, top-N-per-group, slowly-changing dimensions, partition tricks. Writing the shapes by hand turns the unfamiliar into pattern recognition

Related guides