An open-source engine that spreads data processing over a cluster of machines so that big data sets are handled in parallel, in batches or as streams. In Azure you can use it through Microsoft Fabric or Azure Databricks.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Apache Spark in context, with comparison tables and the common traps.
Terms in this definition
- OVER
Gives a T-SQL window function its window: PARTITION BY, ORDER BY and, if wanted, a ROWS or RANGE frame. Rankings and running totals can then be worked out while every row is kept.
- Microsoft Fabric
Analytics platform from Microsoft delivered as SaaS on top of OneLake, offering capabilities including Direct Lake and shortcuts.
- Databricks
Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
Related terms
- ANSI mode
A stricter Spark SQL behaviour, enabled by default since Apache Spark 4.0 (so also in Fabric Runtime 2.0), where a bad cast or an integer overflow fails the query rather than quietly producing null. Functions such as
try_cast,try_addandtry_dividekeep the old null result, while switching ANSI mode off should be treated as a temporary workaround. - Databricks Runtime
The software image, Apache Spark among its core parts, that Azure Databricks installs on compute. Long-term support (LTS) releases are the advised choice when job compute runs production work.
- Distributed processing
Dividing work between several computers, or nodes, in a cluster, with each one handling a share of the data at the same time, as Apache Spark does. This makes it possible to process huge amounts of data.
- Runtime 2.0
The Fabric Spark runtime, now generally available and the recommended choice for production, built on Apache Spark 4.1 with Delta Lake 4.2, Python 3.13, Scala 2.13 and Java 21. Its predecessor, Runtime 1.3 on Spark 3.5, reached end of support on 30 September 2026 and moved into Long Term Support from 1 October 2026 until March 2027.
- SDP
Spark Declarative Pipelines: open-source Apache Spark tooling to declare batch or streaming data flows in Python or SQL. Databricks extends it as Lakeflow Spark Declarative Pipelines, adding
AUTO CDC, expectations and a queryable event log. - Spark Structured Streaming
Apache Spark's engine for near real-time streams, found in both Azure Databricks and Microsoft Fabric. Batch and streaming code share one set of APIs, incoming data is handled in increments, and it pairs naturally with Delta Lake.
- Spark UI
Classic compute's Apache Spark web console. Databricks advises checking the jobs timeline first, then the slowest stage, skew or spill, I/O and lastly the SQL DAG. Serverless and SQL warehouses offer a query profile in its place.