A classic compute option where you give a minimum and maximum worker count and Azure Databricks adjusts the number of workers within that range to match the load.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Autoscaling in context, with comparison tables and the common traps.
Terms in this definition
- Classic compute
All-purpose, jobs and pipeline compute running in your own Azure subscription, which you set up and manage yourself. Its counterpart is serverless compute, which Azure Databricks runs for you.
- WHERE
Limits a SELECT, UPDATE or DELETE to just the rows meeting a condition. Omit it, and the statement hits every row.
- Worker
A classic compute node hosting a Spark executor that carries out distributed work; with one executor on each worker the two words are used as synonyms. Worker node type sets its instance family, and Spark commands fail if there are no workers.
- COUNT
A DAX function returning how many non-blank numbers, dates or text values a column contains. Boolean columns need COUNTA instead, and for counting the rows of a table COUNTROWS is the better choice.
- Databricks
Analytics platform built on Apache Spark where data is processed in notebooks; Unity Catalog is the governance model it recommends.
- RANGE
Usable only in visual calculations, this DAX function picks a span of rows along an axis counted from the current one, say the previous six. Think of it as a simpler WINDOW, handy for moving totals.
- MATCH
Used in WHERE when querying SQL Graph, it describes how to walk from node to node through edge tables, with patterns written like p1-(f1)->p2.
- ELT
Extract, load, transform: raw data lands in the target system first and is transformed there. See ETL for the opposite order.
Related terms
- Apache Airflow job
A code-first Fabric Data Factory item for orchestrating work as Python DAGs on managed Apache Airflow, offering starter or custom pools, autoscaling and syncing DAGs from Git. It succeeds Workflow Orchestration Manager in Azure Data Factory and sits alongside pipelines as their code-based counterpart.
- Application Gateway v2
Present-day Application Gateway SKU, Standard_v2 or WAF_v2, offering zone redundancy, autoscaling, a static VIP, Key Vault integration and header rewrite. It requires its own subnet, ideally a /24.
- az aks
Azure CLI commands for managing AKS. For example,
az aks install-cliinstalls kubectl, while runningaz aks updateoraz aks nodepool updatewith--enable-cluster-autoscaler --min-count --max-countenables autoscaling. - Azure Container Instances
Way to run container groups without any cluster, billed per second. It offers neither autoscaling nor spreading across zones, and Azure Machine Learning treats it as a legacy v1 deployment target.
- Executor
Runs tasks on a worker node; Azure Databricks places exactly one on each worker. Autoscaling, losing a spot instance or out-of-memory errors can all take one away.
- JIT runner
Just-in-time self-hosted runner for GitHub Actions: made via the REST API, it takes one job and then disappears, as with --ephemeral. Handy for autoscaling and limits how long a compromise can persist.
- Larger runners
Bigger GitHub-hosted machines offering extra processor, memory and storage, plus options such as GPUs, autoscaling, fixed IPs and private connection to Azure networks. They come with the GitHub Team and Enterprise Cloud plans and are charged by the minute in every case.
- Managed DevOps Pools
Lets Microsoft run your custom Azure Pipelines agent pools, hosting the VMs in its own subscription while you bring images, a virtual network and autoscaling rules. Microsoft now advises this service over scale set agents.