How purpose, skills and coding level decide between Dataflow Gen2, a pipeline and a notebook, and where Copy job and Apache Airflow jobs fit.
From Ultra Transcenders DP-700 by Tony Rough (coming December 2026)
The three tools overlap, so the exam tests the deciding factor: whether the job is transformation or orchestration, and whether the team works low-code or code-first. A pipeline usually orchestrates the other two rather than replacing them.
Dataflow Gen2 is the Power Query-based, low-code transformation tool; since April 2026 every new Dataflow Gen2 item is created with CI/CD and Git integration support. Pipelines are the low-code orchestration tool that groups activities into a workflow. Notebooks are the code-first Spark (or Python) tool for complex transformation. Microsoft’s data integration decision guide groups them by purpose:
| Aspect | Dataflow Gen2 | Pipeline | Notebook |
|---|---|---|---|
| Primary purpose | Code-free data preparation and transformation | Low-code orchestration: logical grouping of activities | Code-first data preparation and transformation |
| Skill set | ETL, M, SQL (Power Query) | ETL, SQL, plus whatever its activities run | Spark: Python, Scala, Spark SQL, R |
| Coding level | No code or low code | No code or low code | Code-first |
| Transformation support | High: 300+ transformation functions in Power Query | None of its own; it calls activities that transform | High: native Spark and open-source libraries |
| Sources | 170+ built-in connectors plus custom SDK | All Fabric-compatible sources, depending on the activities used | Hundreds of Spark libraries |
| Typical persona | Data engineer, data integrator, business analyst | Data integrator, business analyst, data engineer | Data scientist, developer, data engineer |
| Runs on its own schedule | Yes | Yes | Yes |
| Can be called from a pipeline | Yes, Dataflow activity | Yes, Invoke pipeline activity | Yes, Notebook activity |
Common trap: Choosing a pipeline to transform data because it has the most activities - a pipeline has no transformation support of its own; it orchestrates Copy, Dataflow, Notebook, stored procedure and script activities that do the transforming.
This note is one section of Ultra Transcenders DP-700: Implementing Data Engineering Solutions Using Microsoft Fabric, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Due on Amazon in December 2026, in Kindle and paperback editions.
About the book · DP-700 terms in the glossary · All DP-700 study notes
How authoring style, output, storage, state and latency decide between eventstreams, Spark structured streaming and eventhouses.
When a KQL database should ingest data, query it through a standard OneLake shortcut, or accelerate the shortcut, and what each costs.
How the five eventstream window types group events in time, how they overlap and how to write them in the SQL operator.
When to reload everything or only changes, and which change-detection method catches inserts, updates and deletes.
How starter, custom, capacity and custom live pools differ in node sizes, start-up time, sizing against the capacity and job admission.
The default, email, random and partial masks, the permissions that add or bypass them, and why masking alone doesn't stop inference.
How skills, data location and transformation type decide between Dataflow Gen2, Spark notebooks, KQL update policies and warehouse T-SQL.