Incremental GitHub issues loader with dlt
Follow dlt's own tutorial to load every issue and pull request of a public repository into DuckDB with a Python pipeline. It pages through the GitHub REST API, keeps an updated_at cursor with dlt.sources.incremental so each later run fetches only what changed, and merges rows on id with the merge write disposition, so a rerun never creates a duplicate.
Every API extract needs pagination, retries that respect the server and state that makes the next run incremental, and dlt handles all 3, including retries on 429s and 5xx errors in its requests helper. The endpoint takes since and sort=updated as parameters, along with state=all, and serves up to 100 records a page. It returns pull requests too, marked by a pull_request key. At 60 requests an hour without a token and 5,000 with one, full pages cap a run at 6,000 or 500,000 records an hour.
Then make it yours by writing the core loop by hand in about 40 lines of requests and DuckDB: follow the next URL in the link header, sort newest first, re-read a 5 minute overlap behind the cursor, and advance the cursor only after the rows are stored. Comparing your loop with dlt's shows what the library does for you and where its defaults would not suit your source.
API ingestion is a common take-home, and reviewers ask what happens if the job dies on page 40, how you avoid double counting and how you stay under a rate limit. After this project you can answer each with a line of code from your own loop or from dlt's.



