What Is Data Observability? 5 Pillars, Examples & Tools (2026)

Data observability is the practice of watching data pipelines and the tables they produce closely enough to know when data is late, incomplete, reshaped or wrong, and to trace each break to its source before anyone downstream notices. It has 5 pillars: freshness, volume, schema, quality and lineage. Its measure is data downtime, the hours data spends broken.

Last updated: Proudly published by: Jeff Wahl16 min read

What is data observability?

Data observability is the practice of watching the tables a pipeline produces, not only the jobs that produce them, so that late, partial, reshaped or implausible data is caught by a check before a person catches it on a dashboard. It borrows the idea from software operations, where teams run services by the metrics, logs and traces those services emit.

For data, the signals are metadata the warehouse and the orchestrator already keep: when each table last changed, how many rows each load added, which columns and types each table has, how the values in each column are spread, and which jobs read and write which tables. Checks that read these signals are cheap, because they query metadata and aggregates instead of rescanning whole tables.

Barr Moses of Monte Carlo named the category in 2019 and set out 5 pillars in a December 2020 essay: freshness, distribution, volume, schema and lineage. The vocabulary has since spread well past 1 vendor, and most guides now call the distribution pillar quality. The goal it describes is a short gap between the moment data breaks and the moment its owner knows, with enough context to find the cause without guessing.

Why a successful pipeline can still deliver bad data

Most data incidents raise no error. Nothing throws an exception when a vendor renames a field, when an upstream team loads a partition twice, or when a mobile release stops sending a column. The job exits 0, the dashboard renders, and the number on it is wrong. An orchestrator that reports every run green has said nothing about whether the output is usable.

Without checks on the output, the person who finds the problem is whoever reads the dashboard and doubts it, often a day later. Every hour in that gap is an hour of decisions made on bad data, so observability is judged by how short it makes the gap, not by how many tests exist.

Prepare for the interview
01 / Open invite
02min.

Know Data Observability the way the interviewer who asks it knows it.

a Data Observability query, the same shape a screen would give you.
The diff against expected. Where ties broke. What you missed.
sandbox
1source → bronze → silver → gold
2 ingest : CDC + Kafka
3 transform : dbt + Airflow
4 serve : Snowflake
5
Execute your solution0.4s avg.
SnowflakeInterview question
Solve a Data Observability problem

Freshness: an hourly table that breaches its threshold

Freshness is the time since a table last received new data, measured against how stale that table is allowed to get. It is the first check to build: it needs nothing but a timestamp, and it catches the widest range of failures, from a stuck scheduler to a late upstream extract to a job that succeeded without writing a row.

Take an hourly table, orders_hourly, that should land a new batch every 60 minutes. Its allowed lag is 120 minutes, so 1 late batch is tolerated and 2 missed batches are a breach. The last load finished at 11:05, which means the allowance runs out at 13:05. When the check runs at 14:00, the lag is 175 minutes, past the allowed threshold, and the table is in breach.

Timeline of the hourly table orders_hourly: loads land at 09:05, 10:05 and 11:05, the 12:05 and 13:05 loads are missed, the 120 minute allowance ends at 13:05 and the 14:00 check measures a 175 minute lag, a breachTimeline of the hourly table orders_hourly: loads land at 09:05, 10:05 and 11:05, the 12:05 and 13:05 loads are missed, the 120 minute allowance ends at 13:05 and the 14:00 check measures a 175 minute lag, a breach

With checks every 15 minutes, the alert fires before 14:00: the first run to see a lag over 120 minutes is the 13:15 run, at 130 minutes. Watch 2 clocks where you can. Load time, a loaded_at column the pipeline stamps, says whether the job ran. Event time, the newest created_at inside the data, says whether the source is current. A job that runs every hour against a source that stopped updating yesterday passes the first check and fails the second.

How to set a freshness threshold

Base the threshold on when the consumer reads the table. A table behind an 8:00 executive dashboard must be complete by 7:30 whatever its refresh cadence, while a table read once a week can tolerate a day of lag. A threshold set without the consumer in mind either misses incidents or pages people for delays nobody would notice.

Then allow for normal variance. Hourly jobs finish a few minutes early or late, so a threshold of exactly 1 refresh interval fires on noise. 2 intervals is a common default for hourly tables, often with a warning at 75% of the allowance so someone can look before the breach.

Choose the check cadence last, because it adds straight to detection time. The dbt documentation on source freshness advises running the check at least twice as often as your tightest SLA, so every 30 minutes for a 1 hour SLA. For orders_hourly that means every 60 minutes or less. At 15 minutes, the alert fires no later than 135 minutes after the last good load: the 120 minute allowance plus 1 check interval.

Store the thresholds as data: 1 row per monitored table with its allowed lag, tier and owner. A single scheduled query then judges every table against its own SLA, and adding a table becomes an insert instead of a deploy.

Freshness lag against each table's SLA at 14:00

payments_15min12 min / 30 minfresh
orders_hourly175 min / 120 minbreach
sessions_hourly95 min / 120 minwarning
customers_daily830 min / 1,560 minfresh

Volume: did all of the data arrive?

Volume monitoring counts what each load delivered and compares it with what similar loads delivered before. It is the most dependable early signal of a partial failure: an extract that hit an API rate limit, a partition that never landed, a filter that quietly began excluding a region. Spikes matter as much as drops, because a load that doubles usually carries duplicates.

A fixed floor such as "more than 0 rows" only catches total outages. A baseline catches the partial ones. In the daily orders series, the latest load delivered 10,412 rows after 2 weeks in which no load fell below 42,000: the job succeeded, the floor passed, and 78% of the orders were missing.

Rows loaded into orders per day, with the short load marked

How to detect a volume anomaly

The simplest baseline is the mean and standard deviation of recent loads, with an alert on any load more than 3 standard deviations away. For the daily orders table, the 14 loads before the one being judged average 47,383 rows with a standard deviation of 2,734. A load of 10,412 rows is -13.5 standard deviations from that mean, far outside the band, while a floor of "more than 0 rows" would have passed it.

Elementary's volume_anomalies test for dbt uses the same shape by default: a 14 day training period of daily row counts and an expected range of 3 standard deviations around the mean.

Respect seasonality. Weekends, month ends and holidays move volume on purpose, so either compare each day with the same weekday in earlier weeks or use a window long enough that normal swings widen the band instead of paging someone every Saturday. Judge only complete buckets: a count taken at noon for a day that is half over will always look like a drop.

Schema: did the shape of the data change?

Schema monitoring watches for columns that appear, disappear, get renamed or change type. Most schema changes are deliberate: someone upstream improved their table without knowing a pipeline depended on the old shape. Nothing fails at the source, so the break surfaces 2 or 3 tables downstream as a failed cast or a column that is suddenly all NULL.

The reactive approach snapshots information_schema.columns once a day and diffs each snapshot with the one before it. A column in today's snapshot and not yesterday's was added, a column in yesterday's and not today's was removed, and a column in both with a different data_type changed type. A rename shows up as 1 removal plus 1 addition.

Rank the findings by damage. A removed column fails every query that selects it. A type change can corrupt results through implicit casts without raising an error. An added column is usually safe, although its downstream owners should hear about it before someone builds on it.

The proactive approach blocks breaking changes before they ship: a schema registry that enforces compatibility rules on event streams, or a data contract that turns a producer's schema into a reviewed interface. Most teams run the reactive diff on every table and keep the proactive guard for the few sources that break most often.

Quality: are the values plausible?

The quality pillar, called distribution in the original 2020 framework, looks inside the columns. A load can arrive on time, with the expected row count and an unchanged schema, and still be wrong: a join key that is NULL for 40% of rows, amounts that turned negative, a currency code that arrived in lowercase and now splits every GROUP BY in 2.

Quality checks come in 2 kinds. Rule checks encode what must always hold: order_id is unique, amount is never negative, currency is one of a known set. Statistical checks learn what normal looks like for a column, its null rate, range, cardinality and distribution, and flag drift from it. Rules catch the failures you can predict. Profiles catch the ones you cannot, and that is what observability adds to classic testing.

Profile each load, not the whole table. A null rate of 8% across a year of data can hide a single day at 40%, and that day is the one you need to see. Store each load's rates as they are measured, and alert when a rate leaves the range of its own history.

Lineage: what depends on what?

Lineage is the map of how data moves: which jobs read which tables, which tables feed which models, and which dashboards, exports and machine learning features sit at the end of each path. The other 4 pillars detect a problem. Lineage explains it, in both directions.

Upstream, it narrows the root cause. When an alert fires on a dashboard table, walk the graph toward the sources and stop at the first unhealthy table, because everything past it's a symptom. Downstream, it sizes the blast radius: before you fix, rerun or change a table, lineage lists everything that will move with it and who owns each piece.

Column-level lineage is the strongest form. It records that fct_orders.customer_id comes from raw.orders.customer_id through a specific join, so a NULL spike can be traced to the exact transformation. dbt builds table lineage from the ref() calls in a project. Outside dbt, the OpenLineage specification is the open standard: each job run emits events that name the job, the run and its input and output datasets, with facets such as the schema attached, and a backend such as Marquez, its reference implementation, assembles the graph.

A graph with missing edges stops the upstream walk short of the cause. Most teams start with the lineage their transformation tool produces and add the ingestion and BI edges over time.

How lineage turns 4 alerts into 1 incident

Real incidents rarely trip a single pillar. Say an upstream team ships a change to the orders source: customer_email is dropped and amount becomes a string. Within hours, raw.orders shows a schema change and a short load, stg_orders shows a NULL spike, and orders_hourly breaches freshness because its job keeps failing on the cast. That is 4 alerts that look unrelated, often in 4 different channels.

Lineage turns them into 1 incident. All 4 sit on the same path out of raw.orders, so the first unhealthy node is the source, the owner is the upstream team, and the fix is 1 coordinated change plus a backfill instead of 4 investigations. A freshness warning on sessions_hourly the same afternoon shares no upstream node with them, so it stays a separate, smaller question.

Lineage path raw.orders to stg_orders to orders_hourly grouped as 1 incident: a schema change and a short load on raw.orders, the root cause, a null spike on stg_orders and a freshness breach on orders_hourly; a warning on sessions_hourly stays separateLineage path raw.orders to stg_orders to orders_hourly grouped as 1 incident: a schema change and a short load on raw.orders, the root cause, a null spike on stg_orders and a freshness breach on orders_hourly; a warning on sessions_hourly stays separate

The same walk works for any alert. Confirm the symptom on the alerting table, then move upstream 1 hop at a time and stop at the first node whose checks fail. Before fixing anything, list what sits downstream of that node and who owns it, so tier 1 consumers hear about the problem from you rather than from their own dashboards. Pause the jobs that would spread bad data, repair the cause, rerun the affected partitions in dependency order with jobs that can rerun without double counting, and finish by adding the check that would have caught it sooner.

Talking through an incident like this is a common prompt in the data engineering system design round, where interviewers ask how you would know a pipeline is broken and how you would find the cause.

How to route data alerts without alert fatigue

The fastest way to kill an observability program is to page people for things that do not matter. Every alert needs an owner and a response sized to the damage: a breach on an executive dashboard table deserves a page, a late staging table deserves a message, and a new column deserves a line in a digest.

Tier tables by what they feed, never by their size or cost. Tier 1 holds executive dashboards, finance exports and customer-facing data: freshness SLAs of 30 to 60 minutes, every check on every load, and a page to the on-call engineer. Analyst-facing marts and recurring reports go in tier 2, with SLAs of a few hours and alerts to the owning team's channel during working hours. Staging tables and experiments make up tier 3; they get freshness and schema checks, and their results are summarized in a daily digest.

A useful alert carries what the first responder needs to start working instead of searching: what broke and by how much, the SLA it broke, the owner, the likely upstream cause, what is affected downstream and a link to the runbook. Review the alert log every month. A monitor that fired 3 times running with nobody acting on it is either miscalibrated or watching something nobody needs, and 1 upstream break that fanned out into a dozen alerts is the case for grouping alerts by their shared upstream node.

Tiers, SLAs and alert routing are what interviewers probe when they ask how you would operate a pipeline you designed. The data pipeline interview questions cover that ground, and the data pipeline practice problems let you design the pipelines these checks watch.

How to measure data observability: data downtime

Observability is worth what it saves, and the most useful single measure is data downtime: the time data spends partial, wrong or missing. Monte Carlo popularized the formula, which multiplies the number of incidents by the average time to detect plus the average time to resolve each one.

Written out, the formula shows where monitoring pays. Engineering discipline lowers the incident count and good runbooks shorten resolution, but detection is the term monitoring changes most, because without it detection waits for a stakeholder to notice. Take 8 incidents a month that each take 5 hours to resolve. Found by a stakeholder after 6 hours, they cost 88 hours of downtime. Found by a monitor within 30 minutes, the same 8 incidents cost 44.

Data downtime for 8 incidents on 1 scale: found by a stakeholder, 48 hours of detection plus 40 of resolution make 88 hours; found by a monitor, 4 hours of detection plus the same 40 make 44, saving 44 hoursData downtime for 8 incidents on 1 scale: found by a stakeholder, 48 hours of detection plus 40 of resolution make 88 hours; found by a monitor, 4 hours of detection plus the same 40 make 44, saving 44 hours

Track the 3 inputs separately, so that when downtime moves you know whether the incident count, detection or resolution moved it, and the next investment goes to the input that moved least.

Data observability vs data quality vs monitoring

The 3 terms overlap and are often used interchangeably, but they answer different questions. Data quality asks whether values meet expectations, and it is tested with explicit rules. Monitoring asks whether known metrics crossed known thresholds. Observability asks whether you can explain any failure from the signals you already collect, including failures nobody predicted.

In practice they nest. Quality rules are 1 input to an observability practice, monitoring is how its signals become alerts, and lineage plus history is what lets you explain an alert instead of only receiving it. A team with a handful of critical tables and well understood failure modes can go far with quality tests alone. A team with hundreds of tables and frequent unexplained breaks needs the rest.

Data observability tools

The market splits 3 ways: open-source building blocks, commercial platforms and the checks built into the warehouses themselves. The open-source side covers rules, anomaly tests and lineage well, and asks you to assemble, host and maintain the pieces. Commercial platforms connect to the warehouse, learn baselines for every table and add column-level lineage and incident workflows.

In a dbt project, much of the work is configuration. Source freshness takes warn_after and error_after thresholds, each a count and a period of minutes, hours or days, and reads the timestamp from loaded_at_field. Since dbt 1.9 the freshness block sits under config:, and dbt 1.10 added loaded_at_query for a custom SQL expression. Tests such as not_null, unique and accepted_values encode the quality rules, and Elementary adds volume, freshness and schema change tests that learn each baseline from the table's own history. These configurations come up often in dbt interview questions.

Outside dbt, Great Expectations and Soda Core express checks as code, OpenLineage with Marquez provides lineage, and DataHub and OpenMetadata collect lineage, ownership and test results in a catalog.

The warehouses now ship a first layer of their own. Snowflake's system data metric functions include FRESHNESS, which returns the seconds since a timestamp column's latest value or the table's last modification, plus ROW_COUNT, NULL_COUNT and SCHEMA_CHANGE_COUNT. Databricks data quality monitoring, formerly Lakehouse Monitoring, models each table's commit history and marks a table stale when a commit is unusually late. On Google Cloud, Knowledge Catalog, named Dataplex Universal Catalog until April 2026, runs data quality scans on BigQuery tables. Each covers 1 platform, which is enough when all of your data lives there.

The commercial platforms include Monte Carlo, Bigeye, Anomalo, Sifflet and Acceldata, and the field keeps consolidating: Datadog acquired Metaplane in 2025 to put data observability beside its infrastructure monitoring. Whatever you choose, test how it handles lineage across your actual stack, since root cause analysis is only as good as the graph behind it.

Build vs buy: when a data observability platform earns its cost

Most teams should not start by buying. Freshness and volume checks for the tier 1 tables are a few hundred lines of SQL and a scheduler, and writing them teaches the team what normal looks like for its own data. If the transformations already live in dbt, Elementary is the cheapest step up: anomaly tests and a report with no new infrastructure. If all of the data sits in 1 warehouse, its native checks are a second cheap step.

Buy when the numbers say so: when there are more tables than anyone can tune thresholds for, when incidents regularly cross systems dbt cannot see, or when the hours spent chasing alerts cost more than a license. A platform adds learned baselines on every table from the first day and lineage that crosses tools, so evaluate it against your worst incident from last quarter, not a demo dataset.

How to roll out data observability in 30 days

  1. 01

    Week 1: tier the tables and cover freshness

    List the tables behind dashboards, exports and models, give each a tier and an owning team, and put a freshness SLA on every tier 1 and tier 2 table.

  2. 02

    Week 2: add volume baselines

    Record row counts per load for every tier 1 and tier 2 table, and alert on any complete load more than 3 standard deviations from the 14 loads before it.

  3. 03

    Week 3: watch schemas and critical columns

    Snapshot information_schema.columns daily and diff it, then profile null rates and ranges on the columns that feed joins and money.

  4. 04

    Week 4: wire lineage and routing

    Load lineage from dbt or OpenLineage, group alerts by their shared upstream node, and route by tier with a runbook link on every alert.

Data observability FAQ

What is data observability in simple terms?+
Data observability means knowing whether your data is healthy, and why when it is not, from signals your systems already record. It watches whether data arrived on time, whether all of it arrived, whether its structure changed, whether its values look normal, and how it flows from one system to the next.
What are the 5 pillars of data observability?+
Freshness (is the data on time), volume (did all of it arrive), schema (did its structure change), quality (are the values plausible; the original 2020 framework called this pillar distribution) and lineage (what depends on what). The first 4 detect a problem and lineage locates its cause.
What is an example of a data observability check?+
A freshness check on an hourly table. `orders_hourly` should load every 60 minutes and may lag by at most 120. Its last load finished at 11:05, so the allowance ran out at 13:05. A check at 14:00 measures a lag of 175 minutes, which is past the allowed threshold, so the table is in breach and the check alerts its owner.
Is data observability the same as data monitoring?+
No. Monitoring compares known metrics with known thresholds and reports that something crossed a line. Observability keeps the history, distributions and lineage needed to explain why, including failures nobody wrote a check for. Monitoring is 1 part of an observability practice.
How often should a freshness check run?+
At least twice per SLA window, the rule of thumb in the dbt documentation: every 30 minutes for a 1 hour SLA, and every 60 minutes or less for a 120 minute SLA. The interval adds straight to detection time, so a 120 minute SLA checked every 15 minutes alerts no later than 135 minutes after the last good load.
What is data downtime?+
Data downtime is the time data spends partial, wrong or missing. It is estimated as the number of incidents multiplied by the average time to detect plus the average time to resolve each one, which makes it the most direct measure of whether observability work pays off.
What are the best open-source data observability tools?+
dbt tests and dbt source freshness cover rules and freshness inside dbt projects, Elementary adds anomaly tests on top of dbt, Great Expectations and Soda Core express checks as code, and OpenLineage with its reference backend Marquez provides lineage. DataHub and OpenMetadata collect lineage, ownership and test results in a catalog.
How is data observability different from application monitoring like Datadog?+
Application monitoring watches services: latency, errors and resource use. Data observability watches the data those services produce and move: whether tables are fresh, complete, correctly shaped and plausible. A pipeline can pass every infrastructure metric and still deliver wrong data, which is why the 2 rely on different signals even when 1 vendor sells both.
02 / Why practice

The candidate who gets the offer

  1. 01

    Reading a solution is not the same as writing one

    Every engineer who has frozen on a query they had read a dozen times knows the gap. The only preparation that closes it is producing the answer yourself, under time, before the interview does it for you

  2. 02

    76% of hiring managers reject on the coding task, not the resume

    From HackerRank's 2024 Developer Skills Report. Candidates who look strong on paper still fail the live screen if they haven't done timed, executable practice

  3. 03

    System design comes down to the calls you defend out loud

    Ingestion, batch vs streaming, the bronze/silver/gold layers, idempotency, backfill and replay. Sketching the pipeline and naming the failure modes is the signal, not the boxes

Related guides