The software image, Apache Spark among its core parts, that Azure Databricks installs on compute. Long-term support (LTS) releases are the advised choice when job compute runs production work.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Databricks Runtime in context, with comparison tables and the common traps.
Terms in this definition
- Apache Spark
An open-source engine that spreads data processing over a cluster of machines so that big data sets are handled in parallel, in batches or as streams. In Azure you can use it through Microsoft Fabric or Azure Databricks.
- Databricks
Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
- LTS
Applies to some Databricks Runtime versions, which Learn recommends for job compute in production. Feature work on such a version stops after about six months, and for three years afterwards it gets security and stability fixes backported.
- Job compute
Created by a job to run its tasks and removed when the run is over, this classic compute costs less than all-purpose compute. Either a single task uses it or several tasks within one run share it.
- Agents (classic) API
First-generation Foundry Agent Service API, based on threads, messages and runs. It is deprecated, replaced by conversations and responses, and retires on 31 March 2027.
Related terms
- Credential passthrough
Older Premium-tier Azure Databricks feature that let users reach ADLS under their own identity; it has been deprecated since Databricks Runtime 15.0, and Unity Catalog is the recommended replacement.
- Databricks Runtime for Machine Learning
A flavour of Databricks Runtime that ships with popular ML and DL packages ready to use. Choosing it forces the dedicated mode, since compute in standard mode can't run it.
- Dedicated access mode
Assigns classic compute to one person or one group. Some capabilities work only in this mode, among them R, the RDD API, GPU instances and the ML flavour of Databricks Runtime.
- QUALIFY
A SQL clause, available from Databricks Runtime 10.4 LTS, that filters rows by a window function's result with no subquery needed, such as retaining only
ROW_NUMBER() = 1for each key. - TIMESTAMP_NTZ
A type storing date-time values from year down to second where no operation takes time zones into account, available from Databricks Runtime 13.3 LTS. By contrast
TIMESTAMPapplies the session time zone. - UNPIVOT
Turns several columns into rows, with one column carrying the former column names and value columns carrying their data; it is the opposite of
PIVOTand needs Databricks Runtime 12.2 LTS or later. - VARIANT
A type for semi-structured data like JSON whose structure isn't fixed, available from Databricks Runtime 15.4. It is preferred to storing JSON as strings but can't serve as a partition column.