Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
Also called Azure Databricks.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Databricks in context, with comparison tables and the common traps.
Terms in this definition
- Apache Spark
An open-source engine that spreads data processing over a cluster of machines so that big data sets are handled in parallel, in batches or as streams. In Azure you can use it through Microsoft Fabric or Azure Databricks.
- WHERE
Limits a SELECT, UPDATE or DELETE to just the rows meeting a condition. Omit it, and the statement hits every row.
- Unity Catalog
Azure Databricks' governance solution covering both data and AI in one place, with centralised permissions, auditing, data discovery and lineage.
- Governance
Keeping cloud deployments in line with an organisation's rules on technology, security and compliance, helped by Azure Policy, tags and resource locks.
Related terms
- AI/BI dashboard
The current dashboard type in Azure Databricks: AI helps you author it, its data stays governed by Unity Catalog, and publishing adds a companion Genie Agent. The legacy SQL dashboards it replaced can no longer be opened.
- AQE
Adaptive query execution: Spark re-plans a query while it runs, for instance turning a sort-merge join into a broadcast join, merging shuffle partitions, dealing with skew and propagating empty relations. Databricks advises leaving it enabled.
- Audit log system table
Databricks recommends this over exporting diagnostic logs:
system.access.audit, a Public Preview system table holding audit events from the account's workspaces in a region, retained free for 365 days. - Autoscaling
A classic compute option where you give a minimum and maximum worker count and Azure Databricks adjusts the number of workers within that range to match the load.
- Azure Databricks workspace
A deployed Azure Databricks environment where a group of people build and run notebooks, jobs and compute. Accounts often hold several, each attached to a Unity Catalog metastore within its own Azure region.
- Azure HDInsight
Runs managed clusters for open-source big data frameworks, Hadoop, Spark, Hive and Kafka among them. For new analytics work, Microsoft's current guidance centres on Microsoft Fabric and Azure Databricks.
- Classic compute
All-purpose, jobs and pipeline compute running in your own Azure subscription, which you set up and manage yourself. Its counterpart is serverless compute, which Azure Databricks runs for you.
- Continuous job
A Lakeflow job whose continuous trigger launches a fresh run the moment the last one finishes. For keeping a pipeline running non-stop, Databricks prefers this to the pipeline's own Continuous mode.