How authoring style, output, storage, state and latency decide between eventstreams, Spark structured streaming and eventhouses.
From Ultra Transcenders DP-700 by Tony Rough (coming December 2026)
Fabric gives you three engines that can all touch the same event stream, and they overlap enough that the choice is a common source of confusion. Pick by skills, latency target, state and the store the results must end up in.
The engines are often combined: an eventstream brings events in, routes raw events to an eventhouse for sub-second KQL queries, routes a filtered stream to a lakehouse, and hands alert conditions to Fabric Activator (formerly Data Activator). Figure 11.1 compares the three engines side by side.
| Requirement | Eventstream | Spark structured streaming | Eventhouse (KQL) |
|---|---|---|---|
| Authoring style | No-code canvas, optional SQL operator (preview) | PySpark, Scala or Spark SQL code | KQL queries, management commands, update policies |
| Primary output | Routes to eventhouse, lakehouse, Activator, custom endpoint, derived stream | Delta tables in a lakehouse (also Kafka/Event Hubs sinks) | Native KQL tables queried in place |
| Stores data long term | No (retention 1-90 days, default 1 day) | Yes, in Delta tables | Yes (retention default 3,650 days) |
| Stateful processing | Windowed aggregations, joins | Full: aggregations, stream-stream joins, dedupe, custom state | Update policies, materialized views, query-time windowing |
| Processing model | Continuous, event-based processing | Microbatches by default; Real-time Mode on Runtime 2.0 | Near real time with streaming ingestion; up to the batching time (default 5 minutes) with queued ingestion |
| Best fit | Ingest and route events without code | Lakehouse medallion pipelines, complex code-first logic | Interactive analytics, dashboards, time series, log search |
The Fabric decision guide puts it the same way from the storage side: for streaming event data and high-granularity interactive analytics, use an eventhouse; for big data and data engineering over Delta, use a lakehouse.
Spark structured streaming isn’t automatically the lowest-latency option: it runs microbatches by default, and its Real-time Mode needs Fabric Runtime 2.0 and supports only Kafka-compatible sources and sinks or a custom foreach sink, not Delta tables.
Common trap: Choosing an eventstream as the place to keep historical events - an eventstream only buffers events for its retention period (one day by default, 90 days at most) and you can’t explicitly delete them; persist them by routing to an eventhouse or lakehouse.
This note is one section of Ultra Transcenders DP-700: Implementing Data Engineering Solutions Using Microsoft Fabric, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Due on Amazon in December 2026, in Kindle and paperback editions.
About the book · DP-700 terms in the glossary · All DP-700 study notes
How purpose, skills and coding level decide between Dataflow Gen2, a pipeline and a notebook, and where Copy job and Apache Airflow jobs fit.
When a KQL database should ingest data, query it through a standard OneLake shortcut, or accelerate the shortcut, and what each costs.
How the five eventstream window types group events in time, how they overlap and how to write them in the SQL operator.
When to reload everything or only changes, and which change-detection method catches inserts, updates and deletes.
How starter, custom, capacity and custom live pools differ in node sizes, start-up time, sizing against the capacity and job admission.
The default, email, random and partial masks, the permissions that add or bypass them, and why masking alone doesn't stop inference.
How skills, data location and transformation type decide between Dataflow Gen2, Spark notebooks, KQL update policies and warehouse T-SQL.