How data volume, transformation needs, skills and latency decide which Fabric tool should copy data into OneLake.
From Ultra Transcenders DP-600 by Tony Rough (coming December 2026)
When data does have to be copied, Fabric offers several tools that overlap in places. Learn’s data movement decision guide gives a short rule of thumb: Mirroring for gold data that is already processed, Copy job for raw bronze ingestion, Eventstreams for real-time streaming, and pipelines with the Copy activity when complex orchestration or metadata-driven, parameterised ingestion is needed. Figure 7.1 sets the no-ETL options beside the ingestion tools.
| Tool | Best fit | Skills and interface | Transformation | Incremental loading |
|---|---|---|---|---|
| Copy job | Bulk, incremental or CDC replication without building a pipeline | Wizard; no code or low code | Low (type conversion, column mapping) | Built in: watermark-based or CDC, with state tracked for you |
| Pipeline Copy activity | Large-scale, orchestrated, parameterised movement (petabyte scale) | Wizard and canvas; ETL, SQL, JSON | Low | Manual: watermark logic with expressions and control tables |
| Dataflow Gen2 | Low-code transformation and wrangling with 150+ connectors | Power Query; M | Low to high (300+ transformations) | Append or Replace update methods on each destination |
| Notebook (Spark) | Complex logic, unsupported sources, data science | Code: PySpark, Scala, Spark SQL, R | Low to high | Written in code |
| Eventstream | Real-time event ingestion and routing | No-code canvas | Lightweight stream operations | Continuous |
| Load to Tables | One-off conversion of CSV or Parquet files already in the lakehouse | Lakehouse explorer | None | Append or overwrite only |
| File upload | Small local files, no transformation | Lakehouse explorer | None | Not applicable |
For a warehouse destination, T-SQL ingestion (COPY INTO, CTAS, INSERT...SELECT) is also available; it’s covered in “Data warehouses and T-SQL in Fabric”.
| Setting | Options and defaults |
|---|---|
| Root folder | Tables (managed Delta area) or Files (unmanaged area) |
| Table action | Append; Overwrite (replaces both the data and the schema; previous versions remain available for time travel); Upsert (key columns decide whether a row matches) |
| Enable partition | Partition column types: string, integer, boolean or datetime; partition columns can’t overlap with upsert key columns |
| Apply V-Order | Selected by default; clearing it keeps the original Parquet files without V-Order |
| Files destination copy behaviour | Flatten hierarchy, Merge files, Preserve hierarchy or dynamic content |
| Lakehouse table as source | Table (optionally a past Version or Timestamp) or T-SQL Query (preview) through the SQL analytics endpoint |
| Schema-enabled lakehouse | If no schema is given, dbo is used |
Dataflow Gen2 itself is covered in “Power Query, Dataflow Gen2 and the visual query editor”; the ingestion-relevant settings are these.
saveAsTable, as described in “Lakehouses, Delta tables and Spark transformations”.This PySpark reads CSV files from Files with an explicit schema, which Load to Tables can’t do, and appends them to a Delta table.
from pyspark.sql.types import StructType, StructField, IntegerType, StringType, DateType
schema = StructType([
StructField("order_id", IntegerType(), False),
StructField("customer_id", StringType(), True),
StructField("order_date", DateType(), True)
])
df = spark.read.format("csv").option("header", "true").schema(schema).load("Files/raw/orders/")
df.write.format("delta").mode("append").saveAsTable("bronze.orders")Common trap: Setting a Dataflow Gen2 query to load into a Fabric Warehouse with staging turned off - Warehouse destinations require staging (and a fixed schema); staging is off by default only for Lakehouse and other non-warehouse destinations.
Common trap: Choosing Copy activity table action Overwrite to refresh rows while keeping a hand-tuned table schema - Overwrite replaces both the data and the schema; use Upsert with key columns, or Append, when the table definition must stay.
This note is one section of Ultra Transcenders DP-600: Implementing Analytics Solutions Using Microsoft Fabric, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Due on Amazon in December 2026, in Kindle and paperback editions.
About the book · DP-600 terms in the glossary · All DP-600 study notes
How the two Direct Lake flavours differ in table discovery, permission checks, fallback and unsupported cases, and which one to choose.
When each table storage mode fits a semantic model, based on data size, latency, source security and capacity.
What each Fabric workspace role can do across Power BI, data engineering, warehousing and real-time items.
Which items and settings a deployment copies or leaves alone, and how data source and parameter rules point each stage at its own data.
How data type, team skills, write needs and transactions decide between Fabric's lakehouse, warehouse, eventhouse and other stores.
What the Warehouse can do that a lakehouse's read-only SQL analytics endpoint can't, and when to use each.
The three many-to-many scenarios in a semantic model and the bridge-table or relationship design each one needs.