Python Projects for Data Engineers

9 Real Pipelines to Build

The best Python projects for data engineers are pipelines. Each pulls data from a real API or file set, or from a stream, and loads only what changed; it survives failures and can be rerun without duplicates. These 9 are real public projects, each linked to the page that builds it, which is a repository or course and sometimes a guide, in 3 tiers from an incremental GitHub API loader to change data capture from Postgres, and all 9 run on a laptop for $0. Every dossier covers what you build and on what data, what it teaches, how to extend it and the interview signal it sends, and gives the tools it uses along with its level and time.

Last updated: Proudly published by: Jeff Wahl21 min read

What makes a Python project a data engineering project

A data engineering project moves or models data for other people and systems, or serves it to them and keeps it safe. The test is who depends on the output: a table that an analyst queries every morning, a file another job reads, a replica an application trusts. A game, a to-do app or a script that scrapes quotes into a CSV fails that test however clean its Python is, because nothing downstream relies on it being correct and current, let alone complete.

Python does most of its data engineering work at the edges of the warehouse, where SQL cannot reach: calling APIs and reading files and streams, then validating records before they land and running all of it on a schedule. Every project here makes Python the main tool for 1 of those jobs, and each is judged by the properties a reviewer checks in a real pipeline: it loads only what changed, survives a failure halfway through, gives the same result when run twice, and rejects bad data loudly instead of storing it quietly.

Each of the 9 is a real public project with a page you open and build from. Repositories carry star counts as each showed them on September 27, 2026, and the others are free courses or guides. Projects 1 to 3 move data in from an API, a set of files and a feed that revises itself; projects 4 to 6 add tests and a schedule to that work and reconcile its output; projects 7 to 9 end on concurrency and streaming, with change data capture last. If you are still learning the language, start with Python for data engineering, then warm up with the 10 small Dockerised exercises of the danielbeach/data-engineering-practice repository (2.9k stars).

The 9 Python projects at a glance

#ProjectDifficultyTimeCore stackCost
01Incremental GitHub issues loader with dltBeginner1 weekendPython · dlt · DuckDB$0
02Chunked NYC taxi ingestion, extended into a backfillBeginner+1-2 weekendsPython · Pandas · PostgreSQL$0
03A data contract for the USGS earthquake feedBeginner+1 weekendPython · Pydantic · DuckDB$0
04A test suite for pipeline code, on recorded responsesIntermediate1 weekendPython · pytest · VCR.py$0
05Schedules, partitions and backfills with Dagster EssentialsIntermediate6-10 hoursPython · Dagster · DuckDB$0
06Reconcile a copy against its source with reladiffIntermediate1 weekendPython · reladiff · PostgreSQL$0
07Concurrent Hacker News ingestion with asyncio and threadsAdvanced1-2 weekendsPython · asyncio · aiohttp$0
08A Kafka consumer and PyFlink windows on taxi ridesAdvanced1-2 weekendsPython · Kafka · Flink$0
09Change data capture from Postgres in Python with dltAdvanced2 weekendsPython · PostgreSQL · dlt$0

Star counts on each card are as the repository's page showed them on September 27, 2026; the Dagster course and the free-threading guide are not repositories and carry none.

Beginner Python projects: move data in from APIs and files

Tier 1 · Projects 1-3

These 3 projects cover the 3 ways data reaches a pipeline: an API you page through, files too large to open at once, and a feed that keeps revising what it already sent. All 3 end with the same 2 properties, a table you can reload without duplicates and a record of what was rejected and why.

Project 01 · API ingestion

Incremental GitHub issues loader with dlt

Beginner1 weekendLocal & free

Follow dlt's own tutorial to load every issue and pull request of a public repository into DuckDB with a Python pipeline. It pages through the GitHub REST API, keeps an updated_at cursor with dlt.sources.incremental so each later run fetches only what changed, and merges rows on id with the merge write disposition, so a rerun never creates a duplicate.

Every API extract needs pagination, retries that respect the server and state that makes the next run incremental, and dlt handles all 3, including retries on 429s and 5xx errors in its requests helper. The endpoint takes since and sort=updated as parameters, along with state=all, and serves up to 100 records a page. It returns pull requests too, marked by a pull_request key. At 60 requests an hour without a token and 5,000 with one, full pages cap a run at 6,000 or 500,000 records an hour.

Then make it yours by writing the core loop by hand in about 40 lines of requests and DuckDB: follow the next URL in the link header, sort newest first, re-read a 5 minute overlap behind the cursor, and advance the cursor only after the rows are stored. Comparing your loop with dlt's shows what the library does for you and where its defaults would not suit your source.

Interview signal

API ingestion is a common take-home, and reviewers ask what happens if the job dies on page 40, how you avoid double counting and how you stay under a rate limit. After this project you can answer each with a line of code from your own loop or from dlt's.

PythondltDuckDBrequests
Project 02 · File ingestion

Chunked NYC taxi ingestion, extended into a backfill

Beginner+1-2 weekendsLocal & free

Module 1 of the Data Engineering Zoomcamp builds a Python ingestion script that reads a month of yellow taxi trips with pandas in chunks of 100,000 rows and appends them to Postgres running in Docker. The script takes its connection settings as click options and ships as its own Docker image. Extend it into a backfill: 1 command that loads any range of months, 1 month per run, into a table you can rerun without duplicates.

The tutorial's read_csv(..., iterator=True, chunksize=100000) keeps memory flat whatever the file size; the TLC now publishes Parquet, where PyArrow's ParquetFile.iter_batches does the same job. Its first chunk replaces the table and the rest append, which is safe for 1 file and wrong for a year. Delete or swap 1 month's partition before writing it, so rerunning March replaces March and touches nothing else, and a crash at month 7 leaves months 1 to 6 intact.

The data drifts, which makes it a good backfill test. The TLC publishes monthly files with yellow trips back to 2009, usually after a 2 month delay, and added a cbd_congestion_fee column from 2025 data onwards, so a backfill across 2024 and 2025 meets a new column mid-range. Accept an added nullable column and log it, stop with the file and column named when a column disappears or changes type, and send rows that break a rule, such as a drop-off before the pick-up, to a quarantine file with the rule they failed.

Interview signal

Backfills come up in every pipeline design round: how would you reload last March, what happens if the job dies at month 7, what do you do when a file does not fit in memory. This project gives a tested answer to each, with row counts to back it, on a dataset most reviewers have seen.

PythonPandasPostgreSQLDocker
Project 03 · Validation & contracts

A data contract for the USGS earthquake feed

Beginner+1 weekendLocal & free

Load earthquakes from the USGS event service with the loader pattern of project 1, validating each record with a Pydantic model and sending failures to a quarantine table with their error. Then declare what the table promises in an Open Data Contract Standard YAML file (each field's type, which fields are required and unique, and its quality rules) and run datacontract test against the stored data after every load and in CI, so a broken promise fails the run before it reaches a query.

The catalogue revises its own records, which is what makes it worth a contract. Load by updatedafter, never by event time, or every revision after the first load is lost. The service caps a query at 20,000 events and answers anything larger with HTTP 400, so ask its count method first and split any window above the cap. id is the current preferred id and may change over time, while ids lists every id associated with the event, so match a revised event through ids and 1 earthquake stays 1 row.

The contract states what a correct table looks like. id is required and unique, and time and updated are milliseconds since the epoch. status holds automatic or reviewed, with deleted as the one other allowed value, and latitude stays between -90 and 90. A loader that appends revisions instead of upserting fails the uniqueness rule on its second run, so the contract catches the loader's own bugs as well as bad source data.

Interview signal

Data quality and late-arriving updates come up in most loops, usually as how you handle a record that changes after you loaded it and how a consumer would know your table broke. A loader that has handled real revisions and a changing primary id, with a contract that fails CI when it breaks, answers both from experience.

PythonPydanticDuckDBdatacontract-cli
The loader shape projects 1 to 3 share
Sources
Load
Land
Guard
Serve
API
github issues api
Source
taxi trip files
API
usgs event service
Transform
incremental loader
RETRY6BACKOFFexponentialIDEMPOTENCYupsert on the source id
Pandas
chunked file loader
BACKFILL1 month per run, partition replaced
Storage
cursor state
Storage
quarantine
Storage
landing tables
custom
contract test
ERRORFail the run
Consumer
analyst queries

The incremental loader writes rows before it moves its cursor; the file loader replaces 1 month at a time; both send rejected rows to quarantine with the rule they failed, and the contract test runs on every load.

State after the writeThe cursor advances only after the rows land, so a crash re-reads a window instead of skipping one.
1 row per keyUpserts on the source's own id make every rerun and every overlap window harmless.
The contract is a gateA failed contract test stops the run before a consumer queries a broken table.

How to pick your first Python data engineering project

If your situation is
Pick
Why
You write Python but have never built a pipeline
Projects 1 → 3 → 4
An API loader, a validated table under a contract, then the tests for both: 3 weekends that cover the questions most take-home reviews ask
You come from software engineering
Projects 2 → 5 → 7
You already test and package code, so the gaps are files larger than memory and windowed scheduling with backfills, plus concurrency you have measured
The roles you want mention Kafka or CDC, or streaming in general
Projects 1 → 8 → 9 → 6
Incremental state first, then the same ideas under at-least-once delivery, then a reconciliation that proves the replica matches its source
You are aiming at a platform or senior role
Projects 4 → 6 → 9
Senior loops probe tests and reconciliation as well as CDC, and each attaches to a pipeline you built earlier
You have 1 weekend
Project 1
Add the idempotency test from project 4: a loader that proves it is safe to rerun beats 3 that cannot

Intermediate Python projects: test, schedule and reconcile

Tier 2 · Projects 4-6

A loader that works once is a script; these 3 projects make it a pipeline someone else can run. Each attaches to the loaders of projects 1 to 3 rather than starting over, so the portfolio grows into 1 system instead of 9 unrelated repositories.

Project 04 · Pipeline testing

A test suite for pipeline code, on recorded responses

Intermediate1 weekendLocal & free

Write a pytest suite for the loaders of projects 1 and 3 that runs in seconds with no network. VCR.py records each real HTTP exchange once into a cassette file and replays it on every later run, with the token kept out of the file by filter_headers; hand-edited cassettes add the failures you cannot provoke on demand, such as a 403 with retry-after and a 500 followed by a success. An integration test loads a batch into a DuckDB file under pytest's tmp_path and runs the same load twice to assert the table is identical.

Most pipeline bugs sit away from the happy path, in decisions (does this request retry or fail, does the cursor advance) and in properties (a rerun changes nothing, 1 row per key). Add 1 property test with Hypothesis: for any generated list of records with repeated ids and random update times, the upsert leaves exactly 1 row per id and that row is the newest. Then stop the loader between the load and the cursor write; when you run it again, assert the table converges to the same result.

Writing the suite also changes the loaders themselves. Code that takes an HTTP session and a database connection as arguments can be tested; code that builds its own inside a loop cannot.

Interview signal

How do you test a pipeline is a standard question, and most candidates stop at unit tests. Showing an idempotency test and a crash-and-rerun test says you know what breaks pipelines in production.

PythonpytestVCR.pyHypothesisDuckDB
kevin1024/vcrpy~3.0k★DataRecorded GitHub and USGS responses from projects 1 and 3VCR.py documentation
Project 05 · Scheduling & backfills

Schedules, partitions and backfills with Dagster Essentials

Intermediate6-10 hoursLocal & free

Dagster's free course builds a pipeline on NYC taxi trips as Python assets that load DuckDB. The early lessons cover assets and their dependencies along with definitions and code locations, then resources. Later lessons add schedules and sensors, with partitions and backfills between them, and the course ends in a capstone. By the end, the monthly trip files are a partitioned asset: a schedule adds a partition on each run, and 1 backfill reloads any range of months.

The habit it builds is that every run takes an explicit window. Each partition is a unit of data that runs on its own, and a retry or a backfill works the same way, so the scheduled run and a reload of last March are the same code with a different partition key. The resources lesson moves connections out of the asset code into configuration, which is what makes the same assets testable on a laptop and deployable to a server.

Extend it by moving the loader from project 1 or 2 into the same code location and packaging it with a pyproject.toml. Give it a daily or monthly partition, then check that each partition deletes its own rows before writing, so rerunning a partition never duplicates rows. A loader that already takes a window as input moves into any orchestrator with little change.

Interview signal

How would you deploy and schedule this, and how would you rerun a bad day, follows almost every take-home. A partitioned asset with a schedule and a backfill you have actually run answers both in 1 sentence.

PythonDagsterDuckDBPandas
Project 06 · Reconciliation

Reconcile a copy against its source with reladiff

Intermediate1 weekendLocal & free

Write a Python job that runs after each load and proves the copy is complete and current, not only that the loader exited 0. It diffs a Postgres table against its DuckDB copy by key with reladiff, compares each day's row count for an API source with the count the source reports for the same query (the USGS count method is built for it), and checks the newest updated_at against the clock. It records every result as a row (check, table, window, expected, observed, status) and exits non-zero when a check fails.

reladiff, the maintained fork of the archived data-diff, splits both tables into key ranges, compares a hash of each range inside each database and downloads only the rows of the ranges that differ, so diffing millions of rows moves little data when the copies mostly agree. It supports Postgres and DuckDB among more than 12 engines. Judge completeness against the source, never against your own loader's logs.

Revising sources need a policy: reconcile the last few days on every run, because their counts still move as records are added or reviewed, and older days once. Give each freshness check a tolerance for normal lateness, or it fires every day and gets ignored.

Interview signal

How do you know your pipeline is correct is a question about exactly this. An answer that names a reconciliation against the source, with the day it caught a gap, is stronger than any list of tools.

PythonreladiffPostgreSQLDuckDB
erezsh/reladiff~544★DataYour own Postgres tables and their DuckDB copiesreladiff documentation
Prepare for the interview
01 / Open invite
02min.

Know Python data engineering projects the way the interviewer who asks it knows it.

a Python data engineering projects query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1def sessionize(events):
2 sessions = []
3 for e in events:
4 if gap_minutes(e) > 30:
5
Execute your solution0.4s avg.

Why an incremental API loader sorts newest first

An incremental loader that pages through a list sorted by update time has a subtle failure. A record updated while you crawl jumps to the newest end of the list. In ascending order every record behind it slides back 1 slot, so the record at the page boundary lands on a page you have already read and is never loaded, and nothing reports it.

In descending order the records slide forward instead, so the boundary record is read twice, which an upsert on the id merges, and the updated record is now newer than your cursor, so the next run collects it. Solid cells in the figure are records a fetch read; dashed cells are records it did not.

6 rows sorted by update time, 3 per page, with 1 row updated between the page 1 and page 2 fetches: in ascending order the rows shift back and the boundary row is never read; in descending order they shift forward and the boundary row is read twice6 rows sorted by update time, 3 per page, with 1 row updated between the page 1 and page 2 fetches: in ascending order the rows shift back and the boundary row is never read; in descending order they shift forward and the boundary row is read twice

Order the writes so a crash is harmless: store the rows, then advance the cursor, never the other way round. A crash between the 2 means the next run re-reads a window it already stored, which the upsert absorbs. Re-read a 5 minute overlap behind the cursor for the same reason, since timestamps from a busy API are not a perfect watermark. When the server pushes back, honour it: GitHub's rate limit documentation says to wait for retry-after, or until x-ratelimit-reset when x-ratelimit-remaining is 0, and warns that clients which keep calling while limited can be banned.

Advanced Python projects: concurrency, streaming and CDC

Tier 3 · Projects 7-9

These 3 projects are the same ingestion ideas under harder conditions: millions of requests, a stream that never ends and a database log that replays after a crash. Projects 8 and 9 share 1 design, at-least-once delivery into a sink that makes a replay change nothing.

Project 07 · Concurrent ingestion

Concurrent Hacker News ingestion with asyncio and threads

Advanced1-2 weekendsLocal & free

The Python free-threading guide's asyncio example fetches 100 pages of Hacker News stories and their comments with aiohttp. It runs first on 1 event loop and then on a pool of threads that each run their own loop, and it compares the default build of Python with the free-threaded one. Rebuild it and time it, then point it at the official Hacker News API and fetch every item from id 1 to maxitem with a bounded number of requests in flight. Write the items to Parquet in batches; recording each completed id range lets a crash resume where it stopped.

The guide reports 12 stories a second on 1 thread, 35 with threads on the default build and 80 with threads on the free-threaded build, on a 12-core CPU. The numbers differ because fetching waits on the network, which 1 event loop overlaps well, while parsing HTML runs on a CPU, which only more cores speed up. Measure your own crawl the same way before you choose asyncio over threads, or threads over processes.

On the API, bound concurrency with an asyncio.Semaphore even though the API documents no rate limit, run the fetches in a TaskGroup (Python 3.11 and later) but catch errors per item, because the group cancels every sibling task when 1 fails, and checkpoint completed ranges rather than the highest id seen, since tasks finish out of order. Keep items flagged deleted or dead as states rather than dropping them.

Interview signal

"This ingestion takes 10 hours; make it faster" is a common prompt, and the expected answer separates I/O-bound work (asyncio or threads) from CPU-bound work (processes, or now free-threaded Python). Having measured the difference on a crawl of nearly 50 million items makes that answer concrete.

PythonasyncioaiohttpPyArrow
Project 09 · Change data capture

Change data capture from Postgres in Python with dlt

Advanced2 weekendsLocal & free

Run Postgres in Docker with logical replication switched on, then use dlt's pg_replication source to take an initial snapshot of a table and stream every change after it, deletes included, into DuckDB. The source creates the replication slot and publication it needs and reads changes from the server's built-in pgoutput plugin through psycopg2's LogicalReplicationConnection.

Read its helpers.py before you trust it. It keeps the last commit's log sequence number (LSN) in pipeline state and advances the slot with pg_replication_slot_advance only after the load, instead of psycopg2's send_feedback. The order matters because a slot may send recent changes again after a crash, so apply by primary key and skip anything at or below the stored LSN, and the replay changes nothing.

Then break it on purpose. Kill the pipeline mid-stream and restart it. Then compare the replica with the source row for row, which project 6's reladiff does in 1 command. Stop the consumer for a day and watch the server's disk: a slot keeps every WAL segment its consumer still needs, so an abandoned slot fills it.

Interview signal

CDC is a staple of senior pipeline design rounds, and the follow-ups ask about replays and ordering, and about deletes and a slot that falls behind. After this project you answer them from a replica you have broken and repaired, and the CDC pipeline interview questions are the natural drill once it runs.

PythonPostgreSQLdltDuckDBDocker
dlt-hub/verified-sources~124★DataYour own inserts, updates and deletes on Postgres in Dockerdlt pg_replication documentation
At-least-once in, idempotent apply out (projects 8 and 9)
Sources
Deliver
Apply
Store
Prove
Source
ride producer
PostgreSQL
orders db
Kafka
rides topic
CDC
replication slot
Flink
pyflink windows
IDEMPOTENCYupsert by window key
Transform
cdc loader
IDEMPOTENCYapply by key, skip LSN at or below stored
PostgreSQL
window results
Storage
duckdb replica
custom
reladiff check
ERRORFail the run

Kafka offsets and a replication slot's position both trail the last write, so a restart re-sends recent events; the window upsert and the stored LSN make that replay a no-op, and reladiff proves the replica still matches its source.

Replays are normalA consumer restart or a server crash re-sends recent events; the design assumes it happens.
Apply by keyA primary key upsert, or a skip at or below the stored LSN, turns each replay into a no-op.
Prove itAfter a kill-and-restart test, reladiff compares the replica with its source by key range.

The Column Shuffle

> A long-format metrics export gives you one record per reading, each carrying an `id` and a single `amount`, and the same `id` can recur across many records. For each id, spread its amounts across numbered keys `amount_1`, `amount_2`, and so on in the order they appear, and return one dict per id (its `id` plus those keys) with ids in first-appearance order.

Sample input & expected output(2 examples)
Input · example 1
records:[{"id":"A","amount":10},{"id":"A","amount":20},{"id":"B","amount":5}]
Output
[{"id":"A","amount_1":10,"amount_2":20},{"id":"B","amount_1":5}]
Input · example 2
records:[{"id":"X","amount":1},{"id":"X","amount":2},{"id":"X","amount":3}]
Output
[{"id":"X","amount_1":1,"amount_2":2,"amount_3":3}]

What concurrency buys an I/O-bound crawl

The Hacker News API numbers every item, and its largest id read on 2026-09-27 was 49,865,207. At an assumed 100 ms per request, 1 request at a time moves 10 items a second and needs 57.7 days for the whole range. With 50 requests in flight the same crawl moves 500 a second and finishes in 27.7 hours.

Time to fetch 49,865,207 Hacker News items at 100 ms per request on 1 scale: 57.7 days with 1 request in flight, 27.7 hours with 50 in flightTime to fetch 49,865,207 Hacker News items at 100 ms per request on 1 scale: 57.7 days with 1 request in flight, 27.7 hours with 50 in flight

The work is waiting on the network, so concurrency hides latency instead of adding compute, and the gain holds until something else saturates, whether the server or your bandwidth, or else the CPU that parses each response. That last limit is why a crawl that parses HTML stops scaling on 1 event loop and speeds up again with threads on a free-threaded build of Python. Before adding workers, measure which of the 3 you are waiting on.

Common mistakes in Python data engineering projects

Reloading everything on every run is the most common. When a source offers since or updatedafter, a full reload wastes the rate limit and hides the incremental logic the reviewer came to see. Close behind it is saving the cursor before the load commits, which turns any crash into silent data loss, and appending instead of upserting, which turns any rerun into double counting.

Retry logic goes wrong in both directions. Retrying a 401 or a 404 wastes minutes on an error that will never fix itself, and retrying a 429 without reading the server's wait headers keeps calling a server that has asked you to stop. Retry server errors and rate limits with growing waits, fail fast on the rest, and log every retry.

Trusting a stream or a log to deliver each change exactly once is the advanced version of the same mistake. The Postgres documentation on logical decoding says a slot persists its position only at checkpoint, so after a crash it may send recent changes again, and it warns that a slot nobody reads keeps WAL the server cannot remove. A consumer that applies changes by key survives the first; dropping slots you no longer use prevents the second.

Reading whole files into memory works on a sample and fails on the real month. Tests that call the live API fail when the API is slow and pass when your code is wrong. And a cloned tutorial with no changes, or a project that only exists as a notebook, can't be scheduled or rerun, let alone defended, so the work in it is invisible to a hiring manager however good it is.

The core of an incremental API loader in Python

import json
import os
import time
from datetime import datetime, timedelta
from pathlib import Path

import duckdb
import requests

STATE = Path("state.json")
OVERLAP = timedelta(minutes=5)


def get(session, url, params=None, attempts=6):
    for attempt in range(attempts):
        r = session.get(url, params=params, timeout=30)
        if r.status_code == 200:
            return r
        limited = r.status_code in (403, 429)
        if limited and "retry-after" in r.headers:
            time.sleep(int(r.headers["retry-after"]))
        elif limited and r.headers.get("x-ratelimit-remaining") == "0":
            time.sleep(max(int(r.headers["x-ratelimit-reset"]) - time.time(), 0) + 1)
        elif r.status_code == 429 or r.status_code >= 500:
            time.sleep(min(2 ** attempt, 60))
        else:
            r.raise_for_status()  # a 401 or 404 will not fix itself
    raise RuntimeError(f"gave up on {url} after {attempts} attempts")


def load(repo, db="github.duckdb"):
    state = json.loads(STATE.read_text()) if STATE.exists() else {}
    cursor = datetime.fromisoformat(state.get(repo, "1970-01-01T00:00:00+00:00"))
    params = {
        "state": "all",
        "sort": "updated",
        "direction": "desc",  # an update mid-crawl repeats a row, never skips one
        "since": (cursor - OVERLAP).strftime("%Y-%m-%dT%H:%M:%SZ"),
        "per_page": 100,
    }
    url, rows = f"https://api.github.com/repos/{repo}/issues", []
    with requests.Session() as s:
        s.headers["Authorization"] = f"Bearer {os.environ['GITHUB_TOKEN']}"
        while url:
            r = get(s, url, params)
            rows += [(i["id"], i["number"], i["updated_at"], json.dumps(i)) for i in r.json()]
            url, params = r.links.get("next", {}).get("url"), None  # the link keeps the query

    con = duckdb.connect(db)
    con.execute("""CREATE TABLE IF NOT EXISTS issues (
        id BIGINT PRIMARY KEY, number INTEGER, updated_at TIMESTAMPTZ, payload JSON)""")
    con.executemany("""INSERT INTO issues VALUES (?, ?, ?, ?)
        ON CONFLICT DO UPDATE SET number = EXCLUDED.number,
            updated_at = EXCLUDED.updated_at, payload = EXCLUDED.payload""", rows)
    con.close()

    if rows:  # advance the cursor only after the rows are stored
        state[repo] = max(row[2] for row in rows).replace("Z", "+00:00")
        STATE.write_text(json.dumps(state))

The loop project 1 asks you to write by hand: retries that honour GitHub's rate limit headers, a descending sort with a 5 minute overlap, an upsert on id, and a cursor saved only after the load. Set GITHUB_TOKEN and call load("owner/repo").

Where these projects meet the Python interview

The projects give you material; the Python round still tests you on a clock. Its problems are the same skills at a smaller size: parse a paginated response, deduplicate records by key keeping the newest, retry with backoff, stream a file too large to load, group events into windows. Working the Python interview questions for data engineers shows how those skills are asked, and the Python practice problems let you solve them in a browser against tests, the fastest way to make the patterns in these projects automatic.

Several of these projects are also the shape of a real take-home: an API to ingest, a dataset to clean and a README to defend. The data engineer take-home assignment guide covers how those are scored and how much time to spend, so you can treat project 1 or 3 as a timed rehearsal.

How to turn a finished project into interview answers

  1. 01

    Change it before you publish it

    Swap the data source, add the idempotency test or the contract, and note what broke while you did it, because interviewers recognise the popular tutorials and ask what you changed

  2. 02

    Write down the numbers

    Record rows loaded and runtime for a real run, along with how many requests and retries it took and how many rows it quarantined, because a number is the first thing that makes a project answer believable

  3. 03

    Write the failure story

    Pick 1 thing that broke, whether a crash mid-load or a schema change or rate limit, and write down how you noticed it and what caused it, along with the test that now guards against it

  4. 04

    Draw the design on 1 page

    Sketch how data goes from the source through the loader into the destination and put the state and checks beside that path, then mark where each guarantee lives, since that drawing is what a design round will ask you to produce

  5. 05

    Prepare the 3 follow-ups

    Rehearse what happens on a rerun, what changes at 10 times the volume and what happens when the source changes its schema, because interviewers ask those of almost every project

  6. 06

    Say it out loud under time

    Explain the project in 2 minutes, then take questions for 10, since knowing the answers and delivering them to a stranger on a clock are separate skills

Python projects for data engineers FAQ

What Python projects should a data engineer build?+
Build pipelines: projects that move data from a real source into a store other people query, and keep it correct on every rerun. The strongest set covers an incremental API load that pages through results with retries and keeps its state between runs, a backfill of files larger than memory, a validated table under a data contract, a test suite for the pipeline code, and 1 advanced project to finish with. Concurrency works for that last one, and so do streaming and change data capture.
Are Python projects enough to get a data engineering job?+
They get you past the resume screen and give you material for the loop, but the loop still tests SQL and Python coding on their own, with system design as a separate round. Pair 2 or 3 finished projects with timed practice on interview problems, and use the projects as the examples in your design and behavioral answers.
Should I use a library like dlt or write the loader myself?+
Do both, in that order. A library such as dlt gives you pagination and retries on day 1, with incremental state and merges that already work, which is what a job uses. Then write the core loop by hand once, in about 40 lines, because interviewers ask how the cursor and the upsert work and when a request retries, and the answer cannot be that a library did it.
Do I need pandas for a data engineering project?+
No. Most of the work in these projects is HTTP calls and file handling, and the rest is validating records and loading them. The Zoomcamp's ingestion script reads a file with pandas in chunks, which keeps memory flat, but loading a whole file into a DataFrame is the wrong tool once a file is larger than memory. PyArrow record batches and DuckDB queries over Parquet do the same job with less memory, and Pydantic handles validation.
Should I use Airflow or another orchestrator for a Python project?+
Not at first. Build the pipeline as a command that takes a time window and can be rerun safely, then schedule it with cron. Add an orchestrator when you have several dependent jobs, backfills across many windows or retries you want managed for you; Dagster Essentials is a free way to learn one, and a pipeline that already takes a window as input moves into Airflow with little change, as it does into Dagster or Prefect.
How many projects should I put on my resume?+
2 or 3 finished projects beat 6 half-built ones. Choose them so that together they cover ingestion and correctness on rerun, with at least 1 showing a harder skill. Give each a README with the numbers that prove it runs: rows loaded and runtime for a real run, the requests it made and the failures it survived, and say what you changed from the tutorial it started as.
Is a web scraper a good data engineering project?+
Only when the scraping is the smallest part of it. A scraper that collects quotes into a CSV shows parsing. A pipeline that pulls from a source on a schedule and loads only what changed shows data engineering. It also has to get through failures and rate limits, and it counts only if it validates what it stores and can be rerun without duplicates. A documented API is usually a better source for that than HTML.
02 / Why practice

The candidate who gets the offer

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    5 problem shapes cover 80% of data engineer loops

    Parsing and reshaping, sessionization, dedup with tie-breaks, streaming aggregation, top-N-per-group. Writing them by hand turns the unfamiliar into pattern recognition

Related guides