Spark's engine for near real-time data: you write a batch-like query and it processes new data incrementally, recording progress in a checkpoint with exactly-once guarantees. Auto Loader and streaming tables run on it.
Also called Apache Spark Structured Streaming.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Structured Streaming in context, with comparison tables and the common traps.
Terms in this definition
- Checkpoint
The storage location in which a streaming query records its offsets, state, commits and ID. Thanks to it, a restarted query picks up precisely where it left off. Never point two queries at the same one.
- Auto Loader
The
cloudFilesStructured Streaming source, which picks up new files from cloud storage or volumes as they arrive. Its checkpoint records progress so each file is processed exactly once, and it can infer and evolve the schema. - Streaming
Sending model output piece by piece as server-sent events. Only the delivery changes; the answer's content, completeness and cost stay the same.
Related terms
- DataFrame
Spark's rows-and-columns structure for data. In Structured Streaming it grows continually as each event adds rows, and queries run against it.
- foreachBatch
A Structured Streaming output option, called through
writeStream.foreachBatch, that hands every micro-batch to your own batch code, which suitsMERGEupserts and destinations lacking a streaming writer. Delivery is at least once, so the code should be safe to repeat. - Incremental batch
An approach that runs Structured Streaming with the
availableNowtrigger: it picks up only records added since the previous run and then shuts down, giving scheduled jobs streaming-style progress tracking for the cost of a batch. - Native execution engine
A vectorised engine written in C++, based on Velox and Apache Gluten, that speeds up Fabric Spark by running supported operators natively, notably for Parquet, Delta and CSV. Anything it can't handle, such as JSON, XML or structured streaming, goes back to the JVM engine; you switch it on with spark.native.enabled.
- Output mode
Decides which rows a stateful Structured Streaming query writes out at each trigger: append, the default, for finalised rows only; update for rows that changed, which Delta sinks don't support; or complete for the entire result.
- Real-time Mode
A Structured Streaming mode in Spark, set with Trigger.RealTime on Fabric Runtime 2.0, where tasks stay running and handle each record on arrival rather than working in microbatches. Its output mode must be update, and its sources and sinks are limited to Kafka-compatible ones or a foreach sink, so files and Delta tables can't be used.
- Standard connectors
Part of Lakeflow Connect: Auto Loader for object storage, plus SFTP and message buses (Kafka, Pulsar, Pub/Sub). Choose them over managed connectors when you need more control, working in Databricks SQL, Structured Streaming or pipelines.
- Stateful operation
Any Structured Streaming operation that must hold intermediate data between batches (stream-stream joins, deduplication or aggregations over windows, for instance). That data sits in a state store kept with the checkpoint, and TTL, watermarks and time-range conditions keep it from growing endlessly.