Columnar file format, open source and aimed at analytics. Event Hubs Capture cannot produce it natively.
Read more: Microsoft Learn
In the Ultra Transcenders books
AZ-305DP-900DP-750DP-600DP-700
Each book explains Parquet in context, with comparison tables and the common traps.
Terms in this definition
- FORMAT
A DAX function that turns a value into text according to a format string, for instance "MMMM" to show a month's name. Since the output is text, numeric operations can't use it, and dynamic format strings were introduced to get around that.
- Event Hubs Capture
Automatically saves Event Hubs data as Avro files in ADLS Gen2 or Blob Storage, cutting a file whenever a size or time window is reached.
Related terms
- CETAS
Writes the result of a query out to external storage as files (CSV or Parquet, say), in parallel, and creates an external table over them, all in one T-SQL statement.
- Deletion vectors
Lets Delta Lake and Iceberg tables flag removed or changed rows through metadata rather than rewriting complete Parquet files, a soft delete that speeds up deletes, updates and merges. The physical rewrite happens later, when
OPTIMIZEor aREORG TABLEpurge runs. - Direct copy
Loading a Fabric warehouse by having the Copy activity call the COPY statement straight against Parquet or delimited text files held in Azure Blob Storage or ADLS Gen2, with no intermediate step. If the format or settings aren't supported, staged copy has to be used instead.
- Direct Lake
Mode for Fabric semantic models that loads Delta or Parquet files from OneLake straight into VertiPaq, giving almost import-level performance without duplicating data.
- Disk cache
Keeps copies of remote Parquet files, Delta tables among them, on the local SSDs of worker nodes so the same data is read faster next time. Azure Databricks handles it automatically: no code is required and it invalidates itself when files change, unlike the Apache Spark cache.
- ERRORFILE
Gives COPY INTO somewhere to put rejects: a folder alongside the source files where Fabric saves failed rows to row.csv and the reasons to error.jsonl. Paired with MAXERRORS, a load still counts as successful while rejections stay below that limit, though Parquet loads ignore the option.
- Fabric mirroring
Keeps a live copy of a source system such as Azure Cosmos DB or Azure SQL Database in OneLake, stored as Delta Parquet, so it can be analysed straight away with no ETL to write.
- Fast copy
A Dataflow Gen2 setting that hands large loads to the Copy activity engine used by pipelines, for instance CSV or Parquet files of at least 100 MB, or database sources of 5 million rows or more. With Require fast copy, the refresh fails rather than reverting to the normal engine.