A classic compute node hosting a Spark executor that carries out distributed work; with one executor on each worker the two words are used as synonyms. Worker node type sets its instance family, and Spark commands fail if there are no workers.
Also called worker node.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Worker in context, with comparison tables and the common traps.
Terms in this definition
- Classic compute
All-purpose, jobs and pipeline compute running in your own Azure subscription, which you set up and manage yourself. Its counterpart is serverless compute, which Azure Databricks runs for you.
- Executor
Runs tasks on a worker node; Azure Databricks places exactly one on each worker. Autoscaling, losing a spot instance or out-of-memory errors can all take one away.
Related terms
- Autoscaling
A classic compute option where you give a minimum and maximum worker count and Azure Databricks adjusts the number of workers within that range to match the load.
- Azure Queue Storage
A storage service that keeps very many small messages, each no larger than 64 KB, so parts of an application can pass work to one another asynchronously, such as a website queuing tasks for a background worker.
- Disk cache
Keeps copies of remote Parquet files, Delta tables among them, on the local SSDs of worker nodes so the same data is read faster next time. Azure Databricks handles it automatically: no code is required and it invalidates itself when files change, unlike the Apache Spark cache.
- Origin
Host behind Front Door that serves the content, such as an App Service app, storage, a load balancer or any custom host given by FQDN. An App Service app counts as one origin regardless of its worker count.
- Spot instances
Discounted Azure VMs that can be taken back at short notice. Classic compute keeps the driver on-demand and uses spot for further nodes; if a spot worker is evicted, a spot replacement is tried before an on-demand one.