NYC taxi trips to Parquet with the Zoomcamp batch module
The batch module of the Data Engineering Zoomcamp. Once Spark is installed, you convert 2020 and 2021 of NYC yellow and green taxi trips from CSV to Parquet under an explicit schema. Revenue and trip counts per hour and pickup zone then come from a groupBy and 2 joins, an outer join of the 2 taxi types followed by a join to the zone lookup. Later lessons run the same jobs on a local standalone cluster, on a Dataproc cluster reading Google Cloud Storage, and into BigQuery.
The preparation notebook declares a StructType for each taxi type instead of trusting inference, then writes each month with repartition(4) into its own year and month folder, so the file layout is decided by hand. The lessons on GroupBy and joins then show what each wide step costs as a shuffle, stage by stage.
Make it yours by extending the layout. Write the trips once with partitionBy on pickup year and month instead of path strings, add the 2025 files with their new cbd_congestion_fee column, and confirm in explain() that a filter on 1 month lists only that month under PartitionFilters. Then size the output files against the 128 MB read partitions Spark packs its input into by default.
"How would you partition this table?" is one of the most common Spark design questions. You answer it with the query pattern you partitioned for and the file counts and sizes you measured, backed by a plan that shows the pruning, and you do it on a course reviewers already know.



