The Weekly · Thread

Week 4: The Interface

←

Week 4: The Interface

datadriven·

A hospital's lab results arrive as HL7 messages, a text format older than the web. The quality team needs 1 clean row per result.

The brief is at https://datadriven.io/community/week-4. Ask about the brief, the rules, the schema. Talk about what you are seeing in the sample.

The results are in. This thread is now the place for approaches, write-ups, and what beat you.

27

7 comments

Great! With around 3 keys from the JSON logs, the output has 7 columns. This week's challenge is amazing!

3
datadrivenop·· edited

We're counting on you to grab the #1 slot this week!

5

The sanity check reports rows: "Records produced 1 row" as 0.0 / 300.0 for my submission, even though schema: "Rows matching schema" is 300/300 and keys: "Rows with an unrepeated key" is also 300/300, with overall sanity_status: "ok". what does that exactly mean?

3

The check was counting rows per record as they came out of process(). Your handler holds rows and emits them from finalize(), which is a perfectly good way to solve this one, so every record looked like it produced nothing even though all 300 rows were there and correct. That's why schema and keys read 300/300 and the status was ok: those are measured on the rows you actually emitted, and they were fine. Your score was never affected. The tester only ever looked at the rows, not at which method returned them.
finalize() rows now count toward the records that produced them. Your submission reads 300/300 under the fix. The figure will show correctly on your next submission.

2

I also have the same confusion like @pranav_sarwe also :)))
Let me try a different approach in this week competition :)

0

Thanks for the clarification!

0
Kobe_Bryant·· edited

hello DD, could you check and explain why my code score 0 ?

my code when I run on sample

  • Elapsed time: 1m 15s / 10m
  • Estimated full-stream duration at the current pace: 10 minutes
  • Peak memory usage: 19 / 512 MB
  • Rows matching schema: 300 / 300
  • Records produced: 300 / 300 (1 row per record)
  • Rows with unique keys: 300 / 300
0