Transformation Strategy

Concepts covered: paEltVsEtl, paFullVsIncremental, paLateData, paSchemaEvolution, paDeduplication

Where the transformation runs and how much data each run processes are the 2 decisions the interviewer probes most heavily. The answers reveal whether you understand data modeling rather than data movement. ELT Is the Modern Default ETL transforms data before loading it into the warehouse. ELT loads raw data first, then transforms it inside the warehouse using SQL or dbt. The real difference is where raw data lives after ingestion. In ETL, the warehouse holds only processed data, so reprocessing a historical period requires re-extracting from the source. In ELT, the raw data sits in the bronze layer, so reprocessing means re-running the transform on data already stored. Storage on S3 costs roughly $0.023 per GB per month. Re-extracting 1 TB of historical data from a SaaS API costs hours of

About This Interactive Section

This section is part of the Design a Pipeline: Intermediate lesson on DataDriven, a free data engineering interview prep platform. Each section includes explanations, worked examples, and hands-on code challenges that execute in real time. SQL queries run against a live database. Python runs in a sandboxed Docker container. Data modeling problems validate against interactive schema canvases. All content is framed around what data engineering interviewers actually test at companies like Meta, Google, Amazon, Netflix, Stripe, and Databricks.

How DataDriven Lessons Work

DataDriven combines four interview rounds (SQL, Python, Data Modeling, Pipeline Architecture) with adaptive difficulty and spaced repetition. Easy problems get harder as you improve. Weak concepts resurface until you master them. Your readiness score tracks progress across every topic interviewers test. Every lesson section ends with problems you solve by writing and running real code, not by picking multiple-choice answers.