# Moving Day

> Old schema in, new schema out.

Canonical URL: <https://datadriven.io/problems/moving_day>

Domain: Python · Difficulty: hard · Seniority: L5 · Asked in: Google

## Problem

A schema migration is an ordered list of operations run against every stored record: `rename` a field, `add` one with a default, `remove` one, or `cast` a field to `int`, `float`, or `str`. A field's `path` is dot-separated so an operation can reach into a nested record like `addr.country`, and `add` builds any intermediate records that are missing while `rename`, `remove`, and `cast` skip a path that isn't there. Casting is lenient: a value that will not convert, such as `None`, is left as it is.

## Example 1

Input:

```
records = [{"age":"30","tmp":"x","addr":{"zip":"10001"},"name":"Alice"}], operations = [{"op":"rename","path":"name","new_name":"full_name"},{"op":"cast","path":"age","target_type":"int"},{"op":"remove","path":"tmp"},{"op":"add","path":"addr.country","default":"US"}]
```

Output:

```
[{"age":30,"addr":{"zip":"10001","country":"US"},"full_name":"Alice"}]
```

## Example 2

Input:

```
records = [{"price":"19.99","legacy":"drop-me"}], operations = [{"op":"cast","path":"price","target_type":"int"},{"op":"remove","path":"legacy"},{"op":"add","path":"meta.version","default":2},{"op":"cast","path":"missing","target_type":"float"}]
```

Output:

```
[{"meta":{"version":2},"price":19}]
```

## Worked solution and explanation

### What this really is

Strip off the migration costume and this is a tiny interpreter: each operation names a place in the record (a dot-path) and an action (rename, add, remove, cast). The real skill is separating the two, walk the path to the parent container and the final key, then dispatch on the op kind. Anyone can rewrite a flat field. The trap is that 'addr.country' is not a key, it is a route, and only 'add' is allowed to build that route as it walks it. Miss that and 'add' throws a KeyError on the first missing parent; copy the record shallowly and your nested writes quietly poison the caller's input.

> **Trick to Solving**
>
> Walk the path **to the parent of the leaf**, not all the way to the leaf.
> 
> 1. `parts = op['path'].split('.')`; `*parents, leaf = parts`
> 2. For `add`, create missing intermediate dicts via `setdefault({})`
> 3. For `rename` / `remove` / `cast`, no-op if the leaf isn't there

---

### Break down the requirements

#### Step 1: Per-record, per-op iteration

Process records independently and operations in order. Make a deep copy of each record so you can safely mutate it without affecting the input or sibling records.

#### Step 2: Resolve the dot-path

Split `op['path']` on `.` and walk the record dict to the parent container. For destructive ops (rename, remove, cast), if any segment is missing, skip the op silently. For `add`, use `setdefault({})` at each step so missing intermediates are created.

#### Step 3: Apply the four op kinds

Dispatch on `op['op']`. **rename**: `parent[op['new_name']] = parent.pop(leaf)`. **add**: `parent[leaf] = op['default']`. **remove**: `del parent[leaf]`. **cast**: convert `parent[leaf]` via the target type, but wrap it so a value that will not convert (None, a non-numeric string) is left untouched instead of aborting the whole run. Truncating '123.45' to an int means `int(float(x))`, not `int(x)`.

---

### The solution

**Dot-path resolver with op dispatch**

```python
def migrate_records(records: list[dict], operations: list[dict]) -> list[dict]:
    import copy

    def cast_value(value, target_type):
        try:
            if target_type == 'int':
                return int(float(value))
            if target_type == 'float':
                return float(value)
            return str(value)
        except (ValueError, TypeError):
            return value

    out = []
    for record in records:
        rec = copy.deepcopy(record)
        for op in operations:
            *parents, leaf = op['path'].split('.')
            kind = op['op']
            if kind == 'add':
                container = rec
                for part in parents:
                    container = container.setdefault(part, {})
                container[leaf] = op['default']
                continue
            container = rec
            reached = True
            for part in parents:
                if not isinstance(container, dict) or part not in container:
                    reached = False
                    break
                container = container[part]
            if not reached or not isinstance(container, dict) or leaf not in container:
                continue
            if kind == 'rename':
                container[op['new_name']] = container.pop(leaf)
            elif kind == 'remove':
                del container[leaf]
            elif kind == 'cast':
                container[leaf] = cast_value(container[leaf], op['target_type'])
        out.append(rec)
    return out
```

> **Time and Space Complexity**
>
> **Time:** O(r * o * d) where r is records, o is operations, d is dot-path depth. Each op walks at most d steps.
> 
> **Space:** O(r * s) for the deep-copied output where s is the average record size.

> **Casting has to be lenient**
>
> A cast should never be the reason a whole migration dies. Real records are dirty: a numeric column holds a decimal string, a blank shows up as None. `int('123.45')` raises, so truncate through float first, and wrap the conversion so anything that still won't convert is returned unchanged rather than throwing. That single try/except is the difference between a resilient migration and one that aborts on row 4,000,000.

> **Interviewers Watch For**
>
> Strong candidates separate path resolution from op dispatch and treat `add` as the only path-creating op. They reach for `setdefault` to build intermediates lazily, keep `cast` from raising on bad values, and recognize that skipping a missing leaf makes destructive ops idempotent.

> **Common Pitfall**
>
> Mutating the original record. `dict(record)` is a shallow copy, so `record['addr']` is shared between the input and your output. After `add 'addr.country'` you've poisoned the input. Use `copy.deepcopy(record)` once per record.

---

## Common follow-up questions

- What should happen when `rename` targets a path whose parent doesn't exist? Silent no-op vs raise vs log? _(Tests defensive navigation with try/except or existence checks.)_
- How would you support array indexing in dot paths (e.g., 'items.0.name')? _(Tests parsing numeric path segments as list indices.)_
- If a rename runs before a cast on the same field, do you target the old name or the new one? How does the order of operations encode that? _(Tests operation ordering and dependency resolution.)_
- How would you validate the operations list against a schema before touching any records? _(Tests dry-run validation that checks path existence and type compatibility.)_
- At what data volume would you stop running this in Python and push the migration into Spark or Beam? _(Tests awareness of Spark or Beam for large-scale schema migration.)_

## Related

- [All practice problems](https://datadriven.io/problems)
- [Mock interview mode](https://datadriven.io/interview/moving_day)
- [Python Interview Questions](https://datadriven.io/python-interview-questions)
- [Data Engineering Interview Prep Guide](https://datadriven.io/data-engineer-interview-prep)
- [Daily Challenge](https://datadriven.io/daily)

---

Source: DataDriven (https://datadriven.io). DataDriven is the data engineering interview community. Live code execution in SQL, Python, and Spark sandboxes. Every feature is open to every member.