Unity Catalog managed tables can have OPTIMIZE, VACUUM and ANALYZE scheduled and run for them on serverless compute without manual effort. Accounts created on or after 11 November 2024 get this turned on by default.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Predictive optimization in context, with comparison tables and the common traps.
Terms in this definition
- Unity Catalog
Azure Databricks' governance solution covering both data and AI in one place, with centralised permissions, auditing, data discovery and lineage.
- OPTIMIZE
Bin-packs small files into larger ones, or reclusters liquid-clustered tables, when you run this SQL command; for Unity Catalog managed tables, predictive optimization takes care of it automatically.
- VACUUM
Removes data files that a table no longer references once they are older than the retention period, 7 days unless changed. Versions older than that can no longer be reached with time travel.
- Serverless compute
Azure Machine Learning compute provided on demand whenever a job names no compute target. No cluster needs creating or managing, and jobs do not wait in a queue behind one another.
- AGDLP
Nesting pattern: users go into global groups, which go into domain local groups, which receive the permissions. AGUDLP adds universal groups for forests with several domains.
- Get
Key Vault permission on secrets that allows a single secret to be read; App Service Key Vault references need nothing beyond it.
Related terms
- ANALYZE TABLE
The
ANALYZE TABLE ... COMPUTE STATISTICSstatement gathers statistics on a table and its columns so the cost-based optimiser can plan queries better. On Unity Catalog managed tables, predictive optimization runsANALYZEfor you. - Auto-TTL
Rows older than a set number of days, measured from a timestamp column, are removed automatically (
DELETE ROWS ... DAYS AFTER). It covers Unity Catalog managed Delta or Iceberg tables plus pipeline streaming tables, depends on predictive optimization, and runs at no guaranteed time. - Liquid clustering
Organises a table's data around up to four clustering keys set with
CLUSTER BY, as an alternative to both partitioning andZORDER. The keys can be changed without rewriting existing files, andCLUSTER BY AUTOhands key selection to predictive optimization.