Point at the Wide Op
Concepts covered: paShuffleOptimization
When the follow-up lands, is this narrow or wide, the core move is to point at a specific line and name it. In the query above, the filter is narrow, it runs row by row with no data movement. The groupBy is wide: to sum by category, every row for a given category has to be brought together onto one partition, and that gathering is a shuffle across the network. The orderBy at the end is also wide, because a global sort needs to compare across partitions. Naming which line shuffles, by pointing at it, is the junior answer. Say it against your own code Say it plainly against your code: the filter is narrow and free, the groupBy is the shuffle because it regroups rows by category, and the final sort is a second, smaller shuffle. You have turned a SQL answer into a distributed-execution answer,
About This Interactive Section
This section is part of the SQL at Scale lesson on DataDriven, a free data engineering interview prep platform. Each section includes explanations, worked examples, and hands-on code challenges that execute in real time. SQL queries run against a live database. Python runs in a sandboxed Docker container. Data modeling problems validate against interactive schema canvases. All content is framed around what data engineering interviewers actually test at companies like Meta, Google, Amazon, Netflix, Stripe, and Databricks.
How DataDriven Lessons Work
DataDriven combines four interview rounds (SQL, Python, Data Modeling, Pipeline Architecture) with adaptive difficulty and spaced repetition. Easy problems get harder as you improve. Weak concepts resurface until you master them. Your readiness score tracks progress across every topic interviewers test. Every lesson section ends with problems you solve by writing and running real code, not by picking multiple-choice answers.