Moving Day
A hard Python interview practice problem on DataDriven. Write the Python and run it against test cases, with instant feedback.
- Domain
- Python
- Difficulty
- hard
- Seniority
- L5
- Asked in
Problem
A schema migration is an ordered list of operations run against every stored record: rename a field, add one with a default, remove one, or cast a field to int, float, or str. A field's path is dot-separated so an operation can reach into a nested record like addr.country, and add builds any intermediate records that are missing while rename, remove, and cast skip a path that isn't there. Casting is lenient: a value that will not convert, such as None, is left as it is.
Old schema in, new schema out.
Example 1
Input:
records = [{"age":"30","tmp":"x","addr":{"zip":"10001"},"name":"Alice"}], operations = [{"op":"rename","path":"name","new_name":"full_name"},{"op":"cast","path":"age","target_type":"int"},{"op":"remove","path":"tmp"},{"op":"add","path":"addr.country","default":"US"}]Output:
[{"age":30,"addr":{"zip":"10001","country":"US"},"full_name":"Alice"}]Example 2
Input:
records = [{"price":"19.99","legacy":"drop-me"}], operations = [{"op":"cast","path":"price","target_type":"int"},{"op":"remove","path":"legacy"},{"op":"add","path":"meta.version","default":2},{"op":"cast","path":"missing","target_type":"float"}]Output:
[{"meta":{"version":2},"price":19}]Worked solution and explanation
What this really is
Strip off the migration costume and this is a tiny interpreter. Each operation names a place in the record (a dot-path) and an action (rename, add, remove, cast). The real skill is separating the two: walk the path to the parent container and the final key, then dispatch on the op kind. Anyone can rewrite a flat field. The trap is that addr.country is not a key but a route, and only add is allowed to build that route as it walks it. Miss that and add throws a KeyError on the first missing parent. Copy the record shallowly and your nested writes quietly poison the caller's input.
Break down the requirements
Step 1: Per-record, per-op iteration
Process records independently and operations in order. Make a deep copy of each record so you can safely mutate it without affecting the input or sibling records.
Step 2: Resolve the dot-path
Split op['path'] on '.' and walk the record dict to the parent container. For rename, remove and cast, skip the op silently if any segment is missing. For add, call setdefault(part, {}) at each step so missing intermediates are created.
Step 3: Apply the four op kinds
Dispatch on op['op']. For rename, parent[op['new_name']] = parent.pop(leaf) keeps the field inside the same nested record. For add, parent[leaf] = op['default']. For remove, del parent[leaf]. For cast, convert parent[leaf] to the target type, but wrap it so a value that will not convert (None, a non-numeric string) is left untouched instead of aborting the whole run. Truncating '123.45' to an int means int(float(x)), not int(x).
The solution
def migrate_records(records: list[dict], operations: list[dict]) -> list[dict]:
import copy
def cast_value(value, target_type):
try:
if target_type == 'int':
return int(float(value))
if target_type == 'float':
return float(value)
return str(value)
except (ValueError, TypeError):
return value
out = []
for record in records:
rec = copy.deepcopy(record)
for op in operations:
*parents, leaf = op['path'].split('.')
kind = op['op']
if kind == 'add':
container = rec
for part in parents:
container = container.setdefault(part, {})
container[leaf] = op['default']
continue
container = rec
reached = True
for part in parents:
if not isinstance(container, dict) or part not in container:
reached = False
break
container = container[part]
if not reached or not isinstance(container, dict) or leaf not in container:
continue
if kind == 'rename':
container[op['new_name']] = container.pop(leaf)
elif kind == 'remove':
del container[leaf]
elif kind == 'cast':
container[leaf] = cast_value(container[leaf], op['target_type'])
out.append(rec)
return outCommon follow-up questions
- What should happen when `rename` targets a path whose parent doesn't exist? Silent no-op vs raise vs log? (Tests defensive navigation with try/except or existence checks.)
- How would you support array indexing in dot paths (e.g., `items.0.name`)? (Tests parsing numeric path segments as list indices.)
- If a rename runs before a cast on the same field, do you target the old name or the new one? How does the order of operations encode that? (Tests operation ordering and dependency resolution.)
- How would you validate the operations list against a schema before touching any records? (Tests dry-run validation that checks path existence and type compatibility.)
- At what data volume would you stop running this in Python and push the migration into Spark or Beam? (Tests awareness of Spark or Beam for large-scale schema migration.)