The hardest data engineering test I ever took was a Spark job that had been silently dropping 40% of its records for 6 months. Nobody wrote it as an interview question. It was a Tuesday, finance had a number that looked wrong, and the traceback was empty because nothing had actually failed.
That's the job. A 2022 survey of more than 300 data professionals, run by Wakefield Research for Monte Carlo, put the number on it: 40% of working time goes to evaluating or checking data quality, the average organization eats about 61 data incidents a month, and 75% of respondents need 4 or more hours just to notice one.[1] 2 days a week, every week, spent on data somebody upstream handed you.
So here is the news. DataDriven has launched the Weekly: 1 data engineering challenge a week, built out of exactly that kind of data, that the community competes to solve. You get a brief, a small honest sample, and a starter. You write code. Your last submission before the freeze runs once against a full hidden dataset you never see, and it gets scored. Then everyone finds out what was actually in there.
The first one is open now, for 10 days instead of the usual 7. The rest of this piece is why the format works, why dirty data is the right first test, and how to take part without embarrassing yourself.