The DataDriven Weekly

A Data Engineering Challenge Every Week

DataDriven launches the Weekly: 1 community data engineering challenge a week, scored blind on hidden data. The first is stream parsing, open 10 days.

Published: Proudly published by: Jeff Wahl8 min read

What this post covers

01

How Kaggle turned hard problems into a community sport: Kaggle competition history, participation scale, why the format stuck

02

Capture the Flag as practice for working engineers: CTF history, DEF CON CTF, participation, learning value for practitioners

03

The cost of dirty data in production: Data quality cost estimates, time engineers spend on cleaning, surveys

04

Why stream processing is hard: Out-of-order events, late data, schema drift, at-least-once delivery, upstream changes

05

Why data engineers lack a competitive practice format: Existing DE practice options, gap between interview prep and production work

The hardest data engineering test I ever took was a Spark job that had been silently dropping 40% of its records for 6 months. Nobody wrote it as an interview question. It was a Tuesday, finance had a number that looked wrong, and the traceback was empty because nothing had actually failed.

That's the job. A 2022 survey of more than 300 data professionals, run by Wakefield Research for Monte Carlo, put the number on it: 40% of working time goes to evaluating or checking data quality, the average organization eats about 61 data incidents a month, and 75% of respondents need 4 or more hours just to notice one.[1] 2 days a week, every week, spent on data somebody upstream handed you.

So here is the news. DataDriven has launched the Weekly: 1 data engineering challenge a week, built out of exactly that kind of data, that the community competes to solve. You get a brief, a small honest sample, and a starter. You write code. Your last submission before the freeze runs once against a full hidden dataset you never see, and it gets scored. Then everyone finds out what was actually in there.

The first one is open now, for 10 days instead of the usual 7. The rest of this piece is why the format works, why dirty data is the right first test, and how to take part without embarrassing yourself.

Prepare for the interview
01 / Open invite
02min.

Know the patterns before the interviewer asks them.

a system design query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
PayPalInterview question
Solve a problem

Competitions built the data science and security communities. Data Engineering never got one.

Kaggle is the obvious precedent. Its own milestone post puts registration past 18 million.[2] Whatever you think of leaderboard chasing, that is a community that formed around a shared, scored, repeatable problem, and it produced a generation of people who can actually read a confusion matrix.

Security got there first. DEF CON's capture the flag contest dates back to DEF CON 4 in 1996, and since DEF CON 10 the game has been about running custom services, attacking other teams', and patching and protecting your own.[3] When ENISA studied the format in 2021 it had 879 public CTF events to run statistics on and 22 notable ones to dissect in depth.[4] That is a mature sport. It has qualifiers, dynasties, and write-ups people read for fun.

Notice what both formats reward. Kaggle rewards squeezing signal out of data that was never clean. CTF rewards understanding a system well enough to survive contact with someone who wants it to break. Both are the underlying skill dressed up as a game, and both are more honest tests than the interview loops the same industries run.

I've wanted to play something like that for years. I also spend most of my week figuring out why the pipeline that finance depends on is late, so I never did. Most of y'all are in the same spot. That is the gap the Weekly is aimed at: the same shape of competition, built for the work data engineers actually do, sized so a working engineer can do it in a week.

What the weekly Data Engineering challenge actually is

The mechanics are simple on purpose.[6] Every week a new challenge drops. The brief tells you what the source team claims a record looks like, the exact output schema the downstream team needs, a handful of real records with the rows they should become, and a starter that runs on the sample and nothing more.

You write against the sample. You can run on the sample as often as you like, and it checks every cell against the rows you should have produced. You submit as often as you like too; the last accepted submission before the freeze is the one that counts.

At the freeze, that submission runs once over the full hidden dataset. Your score is accuracy. When 2 runs are equally accurate, the one that does less work per record ranks higher, which means being fast never beats being right, and a careful handler that does the right amount of work is never punished for it.

On reveal day you get your score and your rank, every class of damage the dataset held, and which classes beat you. The week's thread opens for approaches and write-ups. The next week has already dropped.

The sample is honest and small. The dataset holds whatever years of production put there. Everything between those 2 sentences is the challenge.

The shape changes from season to season so it stays fresh; a season of the same problem would teach you 1 trick. And agents are welcome. A Copy button puts the whole brief on your clipboard as plain text, and a Work by MCP option hands an agent the same brief, the same sample run, the same check and the same submit as tools. Bring whatever you work with. The data doesn't care.

Why dirty data is the right first test

Every interview loop I've sat on either side of tests the same 3 things: SQL, some Python, a system design whiteboard. I've written about how to prepare for all of it in the data engineer interview prep guide, and I stand by it; interviewing is a skill and you should treat prep like a job. But none of those rounds test the thing you'll do most.

The Monte Carlo survey again: about half of respondents said resolving an incident takes an average of 9 hours once it's found, on top of the 4 or more hours it took to find.[1] Add it up and the report lands on roughly 793 hours per month, per company, spent identifying and fixing data problems.[1] That is the majority of the craft, and nobody scores you on it until a board deck is wrong.

The failure modes are eternal, too. Schema drift. Late-arriving data. Upstream teams breaking contracts without telling you. A rewrite that changed a field's shape and never migrated the old rows back. The tools change every 18 months; these problems don't.

A challenge built from that kind of data measures the skill that actually compounds: reading a messy record, deciding what it should have been, and writing code that keeps deciding correctly on the 900,000th record when the mess gets creative. That is the reps you can't get from a LeetCode medium.

What makes stream parsing hard

The first challenge is stream parsing, and I want to be precise about why that's difficult rather than wave at it.

The Dataflow paper from Google, the one that shaped how every modern streaming system thinks about time, says it plainly: we must stop trying to groom unbounded datasets into finite pools of information that eventually become complete, and instead live under the assumption that we will never know if or when we have seen all of our data, only that new data will arrive and old data may be retracted.[5] Read that twice. It means every guarantee you'd like to have about order and completeness is something you assume, never something you check.

In practice that gives you 3 constraints at once. You process records as they arrive, 1 at a time, with no second pass to fix your mind. You take the temporal guarantees on faith, because in the real world the stream is usually the only feed left and the people who built it are usually gone. And you consume whatever "gifts" the upstream team left in the payloads, because refusing a record is a decision with a cost downstream, and so is accepting it.

I've argued for years that most companies don't need streaming, and I still think batch covers 90% of what gets built. That is exactly why stream parsing is a good test. The engineer who can hold state correctly across a stream, decide what a later record is allowed to change, and keep the output idempotent under duplicates has the batch problem solved as a special case. The reverse is not true.

What I will not do here is describe the dataset. The brief is public, the sample is public, and the point of the exercise is that the sample doesn't show you everything. Go read it on the community page and draw your own conclusions about what a long-lived export probably contains.

Analysts Are Slowing the Store Down

> Analytics queries against the production database are slowing the live application. Move analytics onto its own warehouse fed from the database's change log, while a merchant dashboard shows new orders within fifteen minutes on a path of its own.

+ Source
+ Transform
+ Storage
+ Quality
+ Consumer
+ Queue
Bronze
Silver
Gold
Custom
Pipeline Architecture
Sketch the architecture.

Click or drag a node from the toolbar above. Right-click the canvas for the full menu.

Drag from a node's right port to another node's left port to wire data flow.

How the inaugural week runs

Because it's the first one, the window is 10 days rather than 7, so nobody finds out about it on Thursday and gets 2 evenings. It opened Thursday, September 10, 2026. Submissions freeze Sunday, September 20 at 23:59 UTC. The reveal is Monday, September 21 at 16:00 UTC. From then on the cadence is Monday to Sunday.[6]

Between now and the freeze, the week's discussion thread is for the brief, the rules, the schema, and what you're seeing in the sample. Solutions wait for the reveal; after that the thread is where approaches, write-ups, and the classes that beat you get argued over. That last part is the point of the whole thing. A leaderboard tells you where you landed. A write-up from the person 3 places above you tells you why.

How to take part without wasting the week

Read the brief the way you'd read a handoff doc from a team that just quit. Every sentence the source team wrote about their data is a claim, and claims are where the bodies are.

Run the starter on the sample first. It works on the sample. That is the last time it will work on anything, and watching where it breaks on your own test lines teaches you more about the shape of the problem than the brief does.

Guard the parse. Normalize every field to the contract, not to what the sample happened to look like. Return a null where there is no usable value instead of inventing one. Keep only the state a later record can legitimately change. That's the whole discipline, and it fits in a few dozen lines if you're honest with yourself.

Submit early, then keep submitting. The sanity check on each submission tells you whether your code survives a run, never how it scores, so you lose nothing by putting something in on day 1. If you want more reps on the same muscles between weeks, the practice problems run your Python and SQL against real data too.

Then show up for the reveal. The score is the least interesting thing on that page.

I've been through 3 waves of "data engineering is getting automated away" and I'm still here, still debugging the same categories of problem. Now there's a place to do it on purpose, against other people, once a week. See you in the thread.

References

  1. Monte Carlo, "Data Engineers Spend Two Days Per Week Firefighting Bad Data, Data Quality Survey Says," August 9, 2022. montecarlo.ai
  2. Kaggle, "One for the numbers: Kaggle registration hits the 18 million mark!" kaggle.com
  3. DEF CON, "Capture the Flag History." defcon.org
  4. ENISA, "CTF Events," May 10, 2021. enisa.europa.eu
  5. Akidau et al., "The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing," Proceedings of the VLDB Endowment, Vol. 8, No. 12, 2015. vldb.org
  6. DataDriven, "The Weekly," community challenge page. datadriven.io
data engineering challengeweekly data engineering challengedirty datastream parsingdata engineering communitydata engineering practice
02 / Why practice

Try the actual problems

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes