How expectations validate records in Lakeflow Spark Declarative Pipelines and what each violation action does.
From Ultra Transcenders DP-750 by Tony Rough (coming November 2026)
Expectations are optional clauses on streaming tables, materialized views and temporary views that test every record against a SQL Boolean constraint and record the results. They need the ADVANCED product edition.
Each expectation has three parts: a name (unique within the dataset, used in metrics), a constraint (valid SQL that returns true or false per record) and an action. Figure 8.1 compares the three actions.
| Action | SQL | Python | Effect on invalid records | Metrics |
|---|---|---|---|---|
| warn (default) | CONSTRAINT n EXPECT (cond) |
@dp.expect("n", "cond") |
Written to the target | Pass and fail counts recorded |
| drop | ... EXPECT (cond) ON VIOLATION DROP ROW |
@dp.expect_or_drop("n", "cond") |
Dropped before the write | Dropped count recorded |
| fail | ... EXPECT (cond) ON VIOLATION FAIL UPDATE |
@dp.expect_or_fail("n", "cond") |
The flow’s update fails and the table update rolls back | Not recorded, because the update fails |
CREATE OR REFRESH STREAMING TABLE orders_silver(
CONSTRAINT valid_id EXPECT (order_id IS NOT NULL) ON VIOLATION DROP ROW,
CONSTRAINT valid_amount EXPECT (amount >= 0) ON VIOLATION FAIL UPDATE,
CONSTRAINT recent EXPECT (order_date >= '2020-01-01')
) AS SELECT * FROM STREAM(orders_bronze);Only Python can group expectations with a collective action. expect_all, expect_all_or_drop and expect_all_or_fail each take one dictionary of name-to-constraint pairs; the single forms take a description and a constraint as two strings. Each expectation in the dictionary still reports its own metrics, and the same dictionary can be reused across datasets or loaded from a rules table.
from pyspark import pipelines as dp
valid_orders = {"valid_id": "order_id IS NOT NULL", "valid_qty": "quantity > 0"}
@dp.table
@dp.expect_all_or_drop(valid_orders)
def orders_clean():
return spark.readStream.table("orders_bronze")In a triggered pipeline, a failed expectation fails and rolls back only that flow; other parallel flows keep updating. In a continuous pipeline, the failure stops the flow and all dependent flows, and the pipeline reports why. Either way, someone must fix the code or data before reprocessing; the error message (EXPECTATION_VIOLATION) shows the violated expectation and, for many queries, the offending input record.
AUTO CDC FROM SNAPSHOT.Common trap: Passing several description and constraint strings to
@dp.expect_all_or_drop- theexpect_alldecorators take a single dict of{name: constraint}; the two-argument form belongs toexpect,expect_or_dropandexpect_or_fail.
Common trap: Assuming a failed expectation in one flow always leaves the rest of the pipeline running - that holds for triggered pipelines only; in a continuous pipeline it also stops every dependent flow.
This note is one section of Ultra Transcenders DP-750: Implementing Data Engineering Solutions Using Azure Databricks, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Due on Amazon in November 2026, in Kindle and paperback editions.
About the book · DP-750 terms in the glossary · All DP-750 study notes
What standard (formerly shared) and dedicated (formerly single user) access modes allow, and when each is required.
Who manages the files, what DROP TABLE does to each, and why Databricks recommends managed tables.
How SQL UDF row filters and column masks restrict data per user, and how they differ from dynamic views.
How the two retention properties and VACUUM decide which table versions you can still query or restore.
SCD types 0, 1, 2 and others compared, and when to keep history in a dimension table.
The table-size thresholds for partitioning, partition sizing, and why liquid clustering is usually the better choice.
Job and task notifications, system destinations, duration warnings and how retries affect which alerts are sent.