The Lakeflow pipelines API for change data capture: it merges incoming changes into a streaming table, keeping either just the latest values (SCD 1) or full history (SCD 2), and uses an ordering column to cope with records that arrive late. A Python-only snapshot variant also exists.
Also called formerly APPLY CHANGES, APPLY CHANGES.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains AUTO CDC in context, with comparison tables and the common traps.
Terms in this definition
- API
Short for application programming interface: a contract that client code calls programmatically, for example a web API secured with tokens or the Files, Images or Responses APIs.
- CDC
Change data capture: rather than reloading whole snapshots, you process just the rows a source has added, changed or removed, with each change carrying its kind and an ordering value. Lakeflow pipelines handle such feeds through
AUTO CDC. - Streaming table
Delta table built for incremental loading, declared with
CREATE STREAMING TABLE(Databricks SQL or Lakeflow pipelines). Wrapping sources inSTREAM, as inSTREAM read_files(...), reads them as streams. - VALUES
Returns in DAX the distinct column values, or table rows, still visible after filters are applied, sometimes with an extra blank entry. CALCULATE often takes the result as a table filter.
- SCD
Slowly changing dimension. Type 1 keeps only current attribute values; type 2 retains every past version, bounded by
__START_ATand__END_AT. Pipelines get both viaAUTO CDC(STORED AS SCD TYPE 1orSTORED AS SCD TYPE 2), whileTRACK HISTORY ONnames which columns produce history. - COPE
Company-owned Android Enterprise devices that people may also use privately. A work profile keeps work apart, and Intune manages that profile as well as the device as a whole.
- snapshot
Captures a VM's disk state, power state, settings and, if requested, memory at a moment in time, using one delta disk per virtual disk. Snapshots should not be relied on as backups.
- VARIANT
A type for semi-structured data like JSON whose structure isn't fixed, available from Databricks Runtime 15.4. It is preferred to storing JSON as strings but can't serve as a partition column.
Related terms
- CDF
Change data feed. When switched on for a Delta table it records which rows changed from version to version, exposing extra columns that give the kind of change and the commit version. You can query it with
table_changes(), stream from it, or feed it toAUTO CDC. - SDP
Spark Declarative Pipelines: open-source Apache Spark tooling to declare batch or streaming data flows in Python or SQL. Databricks extends it as Lakeflow Spark Declarative Pipelines, adding
AUTO CDC, expectations and a queryable event log. - SEQUENCE BY
The part of an
AUTO CDCdefinition that names the column, or aSTRUCTof columns, used to put change records in order so that late or out-of-order events are applied correctly. In Python it issequence_by.