Activity in a Data Factory pipeline that moves data from one store to another.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Copy activity in context, with comparison tables and the common traps.
Terms in this definition
- Azure Data Factory
Managed data integration service for ETL and ELT, built from pipelines, copy activities, triggers, mapping data flows and integration runtimes. It works in batches rather than routing individual transactions.
- Pipeline
Groups activities logically so that, together, they move and transform data, either one after another or side by side; found in both Azure Data Factory and Microsoft Fabric.
Related terms
- Data consistency verification
An optional Copy activity check (
validateDataConsistency) that compares row counts when copying tables, or file size, modified date and checksum when copying binary files. The extra checking makes the copy slower. - Degree of copy parallelism
The upper limit on simultaneous read or write threads in a Copy activity, configured separately from intelligent throughput optimisation. With partitioning in use it governs how many partition queries run together, so a value that is too large risks overwhelming the source system.
- Direct copy
Loading a Fabric warehouse by having the Copy activity call the COPY statement straight against Parquet or delimited text files held in Azure Blob Storage or ADLS Gen2, with no intermediate step. If the format or settings aren't supported, staged copy has to be used instead.
- Dynamic range partition
Parallel reading for the Copy activity, achieved by carving an integer or date/datetime column into slices and running one query per slice. The upper and lower values only decide slice width, never which rows are copied, and hand-written SQL needs the ?DfDynamicRangePartitionCondition placeholder.
- Fast copy
A Dataflow Gen2 setting that hands large loads to the Copy activity engine used by pipelines, for instance CSV or Parquet files of at least 100 MB, or database sources of 5 million rows or more. With Require fast copy, the refresh fails rather than reverting to the normal engine.
- Fault tolerance
Lets a Copy activity keep going instead of failing when it meets rows that don't fit the destination, or files that are missing, forbidden or have invalid names. Skipped rows are counted in rowsSkipped and logging of them is optional.
- Intelligent throughput optimization
A Copy activity setting, stored as dataIntegrationUnits, that limits how much CPU, memory and network a copy may consume; it accepts Auto, Standard, Balanced, Maximum or a number from 4 to 256. The units actually used appear as usedDataIntegrationUnits in the output and determine the data movement charge.
- Partition option
Controls whether a Copy activity reads a relational source in parallel. The default, None, sends one query; Physical partitions of table uses the partitions defined on the table; Dynamic range splits reads by ranges of an integer or date/datetime column. Degree of copy parallelism sets how many reads happen at once.