A cache you request explicitly, in memory or on disk, for a table or DataFrame using df.cache(), df.persist() or CACHE TABLE, and release with unpersist. It differs from the automatic disk cache and cannot be used on serverless compute.
Also called Spark cache.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Apache Spark cache in context, with comparison tables and the common traps.
Terms in this definition
- Event
Table in Log Analytics where entries from Windows event logs are kept.
- DataFrame
Spark's rows-and-columns structure for data. In Structured Streaming it grows continually as each event adds rows, and queries run against it.
- Disk cache
Keeps copies of remote Parquet files, Delta tables among them, on the local SSDs of worker nodes so the same data is read faster next time. Azure Databricks handles it automatically: no code is required and it invalidates itself when files change, unlike the Apache Spark cache.
- Serverless compute
Azure Machine Learning compute provided on demand whenever a job names no compute target. No cluster needs creating or managing, and jobs do not wait in a queue behind one another.
Related terms
- LRU
Short for least recently used. When space runs out, both the Apache Spark cache and the disk cache evict the data that has gone longest without being used.