Orchestration: Beginner
What you will be able to do
Why Cron Is Not an Orchestrator
Recognize the canonical failure mode of cron chains and name what an orchestrator adds beyond a schedule.
What Cron Does and Does Not Do
| Capability | Cron | Orchestrator |
|---|---|---|
| Run a command at a fixed time | Yes | Yes |
| Wait for an upstream job to finish before starting | No | Yes |
| Retry a failing command with backoff | No | Yes |
| Show whether last night's run succeeded | No (just a log file) | Yes (a UI) |
| Re-run a failed task without re-running the whole pipeline | No | Yes |
| Backfill historical date ranges (rerun the pipeline for past days, e.g., the last week of October) | No | Yes |
The Failure That Always Comes First
- ▸Slow upstream jobs cause silent stale reads downstream
- ▸A failed job in the middle does not block the rest of the chain
- ▸There is no UI; status lives in log files spread across servers
- ▸Re-running step 3 after a fix means manually re-running steps 4 and 5
- ▸Backfilling last Tuesday means writing a one-off script that mimics the chain
Why Adding Sleeps Does Not Fix the Problem
- Each job runs at a fixed clock time, regardless of upstream state
- Slow upstream produces stale downstream silently
- Failures do not stop later jobs
- Re-running step N requires re-running everything after it by hand
- Each task runs after its declared dependencies succeed
- Slow upstream delays downstream rather than corrupting it
- A failed task halts dependent tasks automatically
- The orchestrator can re-run a failed task and resume the rest
The Mental Shift
The DAG: Tasks, Edges, No Cycles
Identify nodes, edges, and the acyclic property in a DAG, and read a small DAG declaration that encodes a real pipeline.
Vocabulary, Once and Precisely
| Term | Meaning | Concrete Example |
|---|---|---|
| Task (node) | A single unit of work the orchestrator schedules | Run a SQL query, run a Python script, copy a file |
| Dependency (edge) | A rule that says one task waits for another | join_orders runs after extract_orders finishes |
| Directed | Edges point from upstream to downstream | Data and dependency both flow one way |
| Acyclic | No path leads back to its starting node | extract -> clean -> join, never join -> extract |
| DAG | The whole structure: tasks plus their dependencies | The pipeline a daily ETL declares to the orchestrator |
Why Direction Matters
Why Cycles Are Forbidden
- ▸A new task is added that reads from a table another task in the same DAG overwrites
- ▸Two engineers add edges in the same week without seeing each other's changes
- ▸A backfill task is added at the end of the DAG and depends on the start
- ▸A circular reference in the data model leaks into the dependency graph
A First DAG, in Words
The Same DAG, in Code
- Declare every dependency explicitly in the DAG; never rely on clock-based ordering inside one DAG
- Validate the DAG at deploy time so cycles are caught before they reach production
- Keep tasks small enough that a retry is cheap
- Use one giant task that does extract, clean, and aggregate together
- Add an edge that points back into a task earlier in the DAG
- Treat the schedule as the dependency mechanism within a single DAG
An orchestration DAG: tasks are nodes, dependencies are edges, and there are no cycles. The orchestrator runs each task only after its upstream finishes, retries failures, and backfills - things cron cannot do.
What an Orchestrator Does
Distinguish the four responsibilities an orchestrator owns from the work the orchestrator delegates to other systems.
Retries only produce the same answer as a single run when the work is idempotent: running it twice gives the same result as running it once. That property is the subject of Lesson 5 (idempotency and backfill).
Responsibility 1: Scheduling
Responsibility 2: Dependency Resolution
Responsibility 3: Retries
Responsibility 4: Visibility
| Visibility Surface | What It Shows | Why It Matters |
|---|---|---|
| DAG list | Every pipeline registered with the orchestrator and its current state | On-call sees at a glance which pipelines are healthy |
| Run history | Every prior execution of a DAG with timestamps and status | Trends are visible: a job that gets slower week over week |
| Task instance log | Stdout and stderr of a single task on a single run | The first place a debugger goes when a task fails |
| Graph view | The DAG drawn with nodes colored by state | The shape of the failure is visible: which branch broke |
- ▸Scheduling: when a DAG starts
- ▸Dependency resolution: in what order tasks within the DAG run
- ▸Retries: what happens when a task fails transiently
- ▸Visibility: how operators see what ran, what failed, and why
What the Orchestrator Does Not Own
- When a DAG starts
- What order tasks run in
- Retry policy and failure routing
- The UI that shows run state
- The actual transform: SQL, Spark, Python
- Reading from sources and writing to destinations
- Heavy compute and memory
- The data shape itself
An orchestrator without a UI is a black box. The UI is not a nice-to-have; it is the on-call surface that turns failures into something a human can act on.
The Major Orchestrators by Name
Name the three major orchestrators, describe what each emphasizes, and explain the shared model that makes them interchangeable in concept.
Apache Airflow
Dagster
Prefect
| Orchestrator | Origin | Model | Best Fit |
|---|---|---|---|
| Airflow | Airbnb, 2014 | Task-centric, imperative DAG | Large existing deployments, broad integration needs, stable production |
| Dagster | Elementl, 2018 | Asset-first, typed, software-defined | New builds emphasizing data lineage and testability |
| Prefect | Prefect Technologies, 2018 | Flow-centric, Pythonic, dynamic | Teams that want a hybrid cloud-orchestration model and dynamic graphs |
What They Have in Common
- ▸Argo Workflows: Kubernetes-native orchestrator, used heavily in ML and CI/CD
- ▸Luigi: Spotify's predecessor to Airflow, still in legacy deployments
- ▸Mage: newer, lower-code orchestrator aimed at smaller teams
- ▸Temporal: a workflow engine often used for application orchestration rather than data pipelines
- ▸Cloud-native: AWS Step Functions, Google Cloud Composer (managed Airflow), Azure Data Factory
How to Choose Between Them
- The team already runs Airflow at scale
- A specific operator is needed (rare third-party source)
- Stability and community size outweigh API freshness
- Cloud Composer or MWAA is already in the stack
- A new build with no existing orchestrator
- Asset lineage and software-defined data assets matter (Dagster)
- A hybrid cloud-control plane is preferred (Prefect)
- Local testing and typed pipelines are priorities
First DAG: 3 Tasks, 1 Schedule
Build a three-task DAG with one schedule and one retry policy, and explain the order of execution from the dependencies.
Step 1: Name the Tasks
| Task ID | What It Does | Where It Reads From | Where It Writes To |
|---|---|---|---|
| extract_orders | Pulls new orders from Postgres since the last run | production.orders (Postgres) | raw.orders (Snowflake) |
| clean_orders | Standardizes country codes, drops test accounts | raw.orders | stg.orders |
| aggregate_orders | Counts orders by region for the run date | stg.orders | mart.orders_by_region |
Step 2: Declare the Dependencies
Step 3: Write the Airflow Code
Step 4: Run It and Watch the UI
Step 5: Handle a Failure
- ▸Three tasks, two edges, no cycles
- ▸A schedule (2am daily) that triggers the start of the DAG
- ▸A retry policy applied uniformly to every task
- ▸Dependencies enforced by the orchestrator, not by clock time
- ▸Failure isolation: one failed task halts dependents, not unrelated work
> A startup data team has six cron jobs that run nightly: pull from Postgres, pull from Stripe, clean orders, clean payments, join the two, publish a fact table. The chain has been working for a year. Last week the Postgres pull ran two hours long because of a backfill, and the dashboard showed yesterday's numbers because the downstream jobs ran on stale data. The team asks: 'What is the smallest set of changes that would have prevented this?'
What runs, when, in what order, and what happens when something fails
- Category
- Pipeline Architecture
- Difficulty
- beginner
- Duration
- 25 minutes
- Challenges
- 0 hands-on challenges
Topics covered: Why Cron Is Not an Orchestrator, The DAG: Tasks, Edges, No Cycles, What an Orchestrator Does, The Major Orchestrators by Name, First DAG: 3 Tasks, 1 Schedule
Lesson Sections
- Why Cron Is Not an Orchestrator (concepts: paDagOrchestration)
The first scheduled job most engineers ever write is a cron job. Cron is a Unix utility that runs a command at a fixed time. It is small, reliable, and has been part of every Unix system since 1975. For a single command that runs once a day, cron is the right tool. The trouble starts when several commands need to run in a particular order, and especially when the order has to hold even if one of them runs late. Cron does not know about order. Cron knows about clock time. What Cron Does and Does
- The DAG: Tasks, Edges, No Cycles (concepts: paDagOrchestration)
Every modern orchestrator models a pipeline as a directed acyclic graph, abbreviated DAG. The structure is a small mathematical object with three properties. It has nodes (the tasks). It has edges (the dependencies). The edges point in one direction, and they cannot form a loop. Those properties are not stylistic preferences. They are the conditions that make the graph computable: a structure with cycles cannot be scheduled at all, and a structure without direction cannot be ordered. Vocabulary,
- What an Orchestrator Does (concepts: paDagOrchestration)
An orchestrator is the system that owns four responsibilities: deciding when work runs, running it in the right order, retrying it when it fails, and showing what happened. The four are not separate features bolted together. They reinforce each other. A retry is meaningful only if dependencies are tracked. A schedule is operable only if a UI exists to inspect it. Visibility is useful only if failures are recorded as events the system can react to. Every orchestrator that ships sells the same fou
- The Major Orchestrators by Name (concepts: paDagOrchestration)
Three orchestrators dominate modern data engineering: Airflow, Dagster, and Prefect. Each ships the four responsibilities described in the previous section, but they make different choices in the API and the philosophy. Knowing the names matters because production environments have already chosen one (or, more often, are slowly migrating from one to another). Knowing what they have in common matters more, because the choice of tool changes which buttons are pressed, not what the buttons do. Apac
- First DAG: 3 Tasks, 1 Schedule (concepts: paDagOrchestration)
Vocabulary becomes useful when applied. The example below builds a tiny but complete DAG end to end. A retail company wants a daily summary of orders by region. Three tasks chain together: extract orders from Postgres, clean and standardize the rows, aggregate to one row per region per day. The DAG runs once a day at 2am Pacific. Every concept from the previous sections shows up in working code. Step 1: Name the Tasks Step 2: Declare the Dependencies The dependency graph is a chain. Clean reads