Apache Spark's engine for near real-time streams, found in both Azure Databricks and Microsoft Fabric. Batch and streaming code share one set of APIs, incoming data is handled in increments, and it pairs naturally with Delta Lake.
Also called Apache Spark Structured Streaming.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Spark Structured Streaming in context, with comparison tables and the common traps.
Terms in this definition
- Apache Spark
An open-source engine that spreads data processing over a cluster of machines so that big data sets are handled in parallel, in batches or as streams. In Azure you can use it through Microsoft Fabric or Azure Databricks.
- Databricks
Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
- Microsoft Fabric
Analytics platform from Microsoft delivered as SaaS on top of OneLake, offering capabilities including Direct Lake and shortcuts.
- Streaming
Sending model output piece by piece as server-sent events. Only the delivery changes; the answer's content, completeness and cost stay the same.
- Share
In OpenSharing, the container you fill with tables, views, volumes, models and notebooks for recipients. You need
CREATE SHAREon the metastore to make one; add a schema and all its assets, including later ones, go with history. - Set
Secret permission in Key Vault for writing secrets; some older material refers to it as Create.
- Delta Lake
Table format adding transactions to files in the data lake. In Synapse, Spark pools can write these tables, while serverless SQL pools can only query them.
Related terms
- Micro-batch
Rather than one record at a time, Azure Stream Analytics and Spark Structured Streaming process incoming streamed records in small groups as they arrive; each group is a micro-batch.