Home › Glossary › Apache Spark cache

Apache Spark cache

A cache you request explicitly, in memory or on disk, for a table or DataFrame using df.cache(), df.persist() or CACHE TABLE, and release with unpersist. It differs from the automatic disk cache and cannot be used on serverless compute.

Also called Spark cache.

Read more: Microsoft Learn

In the Ultra Transcenders books

DP-750

Each book explains Apache Spark cache in context, with comparison tables and the common traps.

Terms in this definition

Related terms

See Apache Spark cache in the full glossary