Spark writing data to disk because shuffles, joins, sorts or aggregations have exhausted execution memory, which is costly. It is reported in the stage details of the Spark UI.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Spill in context, with comparison tables and the common traps.
Terms in this definition
- Aggregations
Summary queries over a large DirectQuery table can be answered from memory instead of the source thanks to this semantic model feature: it keeps a concealed summary table cached and sends qualifying queries there, while anything needing fine detail still goes back to the source.
- Spark UI
Classic compute's Apache Spark web console. Databricks advises checking the jobs timeline first, then the slowest stage, skew or spill, I/O and lastly the SQL DAG. Serverless and SQL warehouses offer a query profile in its place.
Related terms
- Shuffle
When Spark has to move rows between executors, as joins, sorts and aggregations require. Per stage, the Spark UI lists Shuffle Read and Shuffle Write; heavy shuffling frequently causes spill.
- SQL DAG
Graph showing the physical plan of a Spark query, reached through Associated SQL Query in the Spark UI. Operators report row counts, spill and peak memory; their timings add up task time.