
An independent study guide for Microsoft Certified: Fabric Data Engineer Associate · by Tony Rough
Know how to build, run and tune data engineering solutions in Microsoft Fabric, and why each setting matters.
Due on Amazon in December 2026, in Kindle and paperback editions.
This independent study guide for the Microsoft Certified: Fabric Data Engineer Associate exam distils what DP-700 really expects you to understand into the comparisons, configuration choices and traps that data engineering decisions turn on, with short PySpark, Spark SQL, T-SQL, KQL and pipeline expression examples throughout.
Fifteen chapters, each readable on its own and together covering all three DP-700 skill areas:
Microsoft Fabric changes quickly. This edition reflects Microsoft's documentation as of October 2026 and the skills measured from 19 October 2026, including Fabric Runtime 2.0, NotebookUtils (formerly MSSparkUtils), Fabric Activator (formerly Data Activator), the Monitor hub and workspace monitoring.
This book contains no exam questions. It explains the knowledge the exam expects, so you can answer questions you have never seen and apply the same judgement to real data platforms.
Written by Tony Rough, a cloud architect with more than twenty years in IT infrastructure who holds the Azure Solutions Architect Expert, Azure Administrator and Azure Security Engineer certifications.
Part of the Ultra Transcenders series from Distilled Press. An independent publication, not affiliated with, sponsored by or endorsed by Microsoft Corporation.
Every skill area in Microsoft's DP-700 outline (as of October 19, 2026), and the chapters that cover it.
| Skill area | Weight | Chapters |
|---|---|---|
| Implement and manage an analytics solution | 30–35% | 1, 2, 3, 4, 5, 6 |
| Ingest and transform data | 30–35% | 1, 7, 8, 9, 10, 11, 12 |
| Monitor and optimize an analytics solution | 30–35% | 13, 14, 15 |
Plus an appendix glossary of 400+ terms, each linked to Microsoft Learn, with the same terms explained free online for print readers.



Some sections of the book, free to read online:
How purpose, skills and coding level decide between Dataflow Gen2, a pipeline and a notebook, and where Copy job and Apache Airflow jobs fit.
How authoring style, output, storage, state and latency decide between eventstreams, Spark structured streaming and eventhouses.
When a KQL database should ingest data, query it through a standard OneLake shortcut, or accelerate the shortcut, and what each costs.
How the five eventstream window types group events in time, how they overlap and how to write them in the SQL operator.
When to reload everything or only changes, and which change-detection method catches inserts, updates and deletes.
How starter, custom, capacity and custom live pools differ in node sizes, start-up time, sizing against the capacity and job admission.
The default, email, random and partial masks, the permissions that add or bypass them, and why masking alone doesn't stop inference.
How skills, data location and transformation type decide between Dataflow Gen2, Spark notebooks, KQL update policies and warehouse T-SQL.