FREE STUDY NOTES · DP-600

Choosing a Fabric ingestion tool: Copy job, pipelines, Dataflow Gen2, notebooks or eventstreams

How data volume, transformation needs, skills and latency decide which Fabric tool should copy data into OneLake.

From Ultra Transcenders DP-600 by Tony Rough (coming December 2026)

When data does have to be copied, Fabric offers several tools that overlap in places. Learn’s data movement decision guide gives a short rule of thumb: Mirroring for gold data that is already processed, Copy job for raw bronze ingestion, Eventstreams for real-time streaming, and pipelines with the Copy activity when complex orchestration or metadata-driven, parameterised ingestion is needed. Figure 7.1 sets the no-ETL options beside the ingestion tools.

A decision splits into no-ETL options (OneLake shortcuts, which point to data in place, and mirroring, which keeps a read-only Delta replica) and ingestion tools. The ingestion side lists when to copy and the best fit for Copy job, pipeline Copy activity, Dataflow Gen2, notebooks and eventstreams.
Figure 7.1: Accessing data with no ETL versus ingesting it with a copy tool
Tool Best fit Skills and interface Transformation Incremental loading
Copy job Bulk, incremental or CDC replication without building a pipeline Wizard; no code or low code Low (type conversion, column mapping) Built in: watermark-based or CDC, with state tracked for you
Pipeline Copy activity Large-scale, orchestrated, parameterised movement (petabyte scale) Wizard and canvas; ETL, SQL, JSON Low Manual: watermark logic with expressions and control tables
Dataflow Gen2 Low-code transformation and wrangling with 150+ connectors Power Query; M Low to high (300+ transformations) Append or Replace update methods on each destination
Notebook (Spark) Complex logic, unsupported sources, data science Code: PySpark, Scala, Spark SQL, R Low to high Written in code
Eventstream Real-time event ingestion and routing No-code canvas Lightweight stream operations Continuous
Load to Tables One-off conversion of CSV or Parquet files already in the lakehouse Lakehouse explorer None Append or overwrite only
File upload Small local files, no transformation Lakehouse explorer None Not applicable

For a warehouse destination, T-SQL ingestion (COPY INTO, CTAS, INSERT...SELECT) is also available; it’s covered in “Data warehouses and T-SQL in Fabric”.

Copy job

Pipeline Copy activity into a lakehouse

Setting Options and defaults
Root folder Tables (managed Delta area) or Files (unmanaged area)
Table action Append; Overwrite (replaces both the data and the schema; previous versions remain available for time travel); Upsert (key columns decide whether a row matches)
Enable partition Partition column types: string, integer, boolean or datetime; partition columns can’t overlap with upsert key columns
Apply V-Order Selected by default; clearing it keeps the original Parquet files without V-Order
Files destination copy behaviour Flatten hierarchy, Merge files, Preserve hierarchy or dynamic content
Lakehouse table as source Table (optionally a past Version or Timestamp) or T-SQL Query (preview) through the SQL analytics endpoint
Schema-enabled lakehouse If no schema is given, dbo is used

Dataflow Gen2 destinations

Dataflow Gen2 itself is covered in “Power Query, Dataflow Gen2 and the visual query editor”; the ingestion-relevant settings are these.

Load to Tables, notebooks and eventstreams

This PySpark reads CSV files from Files with an explicit schema, which Load to Tables can’t do, and appends them to a Delta table.

from pyspark.sql.types import StructType, StructField, IntegerType, StringType, DateType

schema = StructType([
    StructField("order_id", IntegerType(), False),
    StructField("customer_id", StringType(), True),
    StructField("order_date", DateType(), True)
])

df = spark.read.format("csv").option("header", "true").schema(schema).load("Files/raw/orders/")
df.write.format("delta").mode("append").saveAsTable("bronze.orders")

Common trap: Setting a Dataflow Gen2 query to load into a Fabric Warehouse with staging turned off - Warehouse destinations require staging (and a fixed schema); staging is off by default only for Lakehouse and other non-warehouse destinations.

Common trap: Choosing Copy activity table action Overwrite to refresh rows while keeping a hand-tuned table schema - Overwrite replaces both the data and the schema; use Upsert with key columns, or Append, when the table definition must stay.

Get the whole book

This note is one section of Ultra Transcenders DP-600: Implementing Analytics Solutions Using Microsoft Fabric, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Due on Amazon in December 2026, in Kindle and paperback editions.

About the book · DP-600 terms in the glossary · All DP-600 study notes

More DP-600 study notes