A vectorised engine written in C++, based on Velox and Apache Gluten, that speeds up Fabric Spark by running supported operators natively, notably for Parquet, Delta and CSV. Anything it can't handle, such as JSON, XML or structured streaming, goes back to the JVM engine; you switch it on with spark.native.enabled.
Also called NEE.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Native execution engine in context, with comparison tables and the common traps.
Terms in this definition
- LAMP
Short for Linux, Apache, MySQL and PHP (Perl and Python also fill the P): a widely used open-source stack for web applications whose database is frequently Azure Database for MySQL.
- Parquet
Columnar file format, open source and aimed at analytics. Event Hubs Capture cannot produce it natively.
- Cluster Shared Volumes
Shared failover cluster storage that all nodes write to simultaneously, typically for Scale-Out File Server or Hyper-V. On SAN-backed storage, ReFS-formatted volumes work in redirected mode.
- JSON
Text-based format for structured data. ARM templates and Cosmos DB documents are written in it, and most Azure REST APIs exchange request and response bodies as application/json.
- XML
A text-based format for data, used for instance in message bodies.
- Structured Streaming
Spark's engine for near real-time data: you write a batch-like query and it processes new data incrementally, recording progress in a checkpoint with exactly-once guarantees. Auto Loader and streaming tables run on it.
- JVM
The Java Virtual Machine hosting the Spark driver and executors. When tasks share a job cluster they use one driver JVM, so libraries already loaded and Scala singleton state persist from one task to the next.
- SWITCH
Checks one expression against several candidate values, returning whichever result pairs with the match (or a fallback otherwise). Writing SWITCH ( TRUE (), ... ) avoids nested IF chains: whichever condition is first true wins.