Organises a table's data around up to four clustering keys set with CLUSTER BY, as an alternative to both partitioning and ZORDER. The keys can be changed without rewriting existing files, and CLUSTER BY AUTO hands key selection to predictive optimization.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Liquid clustering in context, with comparison tables and the common traps.
Terms in this definition
- Event
Table in Log Analytics where entries from Windows event logs are kept.
- Set
Secret permission in Key Vault for writing secrets; some older material refers to it as Create.
- Partitioning
Splits a table into directories, one per value of the column or columns named in
PARTITIONED BY. Learn favours liquid clustering in most cases and keeps this technique for very large tables. - Index field attributes
Settings applied to each field in an Azure AI Search index:
searchablefor full text,retrievableto return it,filterablefor exact-match$filter,sortable,facetablefor counts, andkeyfor the unique document ID. - Predictive optimization
Unity Catalog managed tables can have
OPTIMIZE,VACUUMandANALYZEscheduled and run for them on serverless compute without manual effort. Accounts created on or after 11 November 2024 get this turned on by default.
Related terms
- Data clustering
A Fabric Data Warehouse preview capability that, as data is loaded, keeps rows with similar values in up to four chosen columns physically together, letting filtered queries avoid reading irrelevant files. You set it once with
WITH (CLUSTER BY (...))inCREATE TABLEor CTAS and cannot alter it afterwards; lakehouse Delta liquid clustering is a different feature. - Data skipping
Avoiding reads of files that cannot hold matching rows, using column statistics such as min, max and null counts gathered per file at write time. Liquid clustering and Z-ordering improve how much gets skipped.
- Fast optimize
With spark.microsoft.delta.optimize.fast.enabled switched on, Fabric Spark's OPTIMIZE skips compacting small-file groups unlikely to hit the target size, so fewer files are rewritten. The setting has no bearing on Z-Order or liquid clustering.
- Z-ordering
Older technique:
OPTIMIZE ... ZORDER BY (cols)clusters related values into the same files so more data can be skipped. Incompatible with liquid clustering, which is preferred for new tables.