The Straggler Task
Concepts covered: paDataSkew
Skew announces itself with one unmistakable shape: the straggler. A stage launches 200 tasks. In the first half minute, 199 of them finish. The last one keeps running, for 10 minutes, for 45, sometimes for hours, while the stage progress bar sits one tick from complete. Nothing errors. Nothing retries. The job is not stuck in any way a health check can see; it is simply one task working through a partition many times bigger than everyone else's. If you spend time around production Spark, you will see this shape so often it becomes reflex to name it before you have opened a single tab. The reason one task can hold a whole stage is the barrier you met in the shuffle lesson. A stage does not end until every one of its tasks ends, because the next stage needs all of the regrouped data in place
About This Interactive Section
This section is part of the Data Skew: Beginner lesson on DataDriven, a free data engineering interview prep platform. Each section includes explanations, worked examples, and hands-on code challenges that execute in real time. SQL queries run against a live database. Python runs in a sandboxed Docker container. Data modeling problems validate against interactive schema canvases. All content is framed around what data engineering interviewers actually test at companies like Meta, Google, Amazon, Netflix, Stripe, and Databricks.
How DataDriven Lessons Work
DataDriven combines four interview rounds (SQL, Python, Data Modeling, Pipeline Architecture) with adaptive difficulty and spaced repetition. Easy problems get harder as you improve. Weak concepts resurface until you master them. Your readiness score tracks progress across every topic interviewers test. Every lesson section ends with problems you solve by writing and running real code, not by picking multiple-choice answers.