The collective term in Fabric for what executes notebooks and Spark job definitions: runtime version, pools, node sizes and Spark properties. Settings can be made per capacity, per workspace, per environment or per session, and through Spark Compute settings capacity admins govern choices like letting workspaces customise their pools.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Spark compute in context, with comparison tables and the common traps.
Terms in this definition
- Job
A sequence of steps executed together on one agent or runner, or on the server for agentless work. While running, each occupies one of your parallel jobs.
- Capacity
A reserved block of compute for a tenant whose size is fixed by its SKU; Fabric F SKUs express it in capacity units (CUs). Workspaces assigned to it unlock features like Copilot, and from F64 upwards people with free licences can view Power BI content.
- Workspace
Teams in Power BI and Microsoft Fabric collaborate in this folder-style container, which groups items such as reports, semantic models and lakehouses, controls who can access them and is assigned a capacity.
- Azure Machine Learning environment
Versioned asset that pairs a Docker image with a pip or conda specification, so every job or deployment using it gets identical dependencies. You point to it by name plus a version number, or by name with @latest.
- Govern
The part of the Cloud Adoption Framework concerned with keeping cloud use under control by setting guardrails made of policies, procedures and tools. It runs continuously across the estate alongside Secure and Manage.
- LIKE
Compares strings with a pattern that can contain the % and _ wildcards. Because it only understands character patterns, searching big volumes of text this way is much slower than using full-text search.
Related terms
- Apache Spark pool
Spark compute in Azure Synapse used to build and update Delta Lake tables, run machine learning and stream changes from the Cosmos DB change feed.
- Capacity pools
Custom Spark pools defined by a capacity admin in the Spark Compute settings of a capacity, which then show up as choices for workspaces and environments and take roughly two to three minutes to start. From the same place the admin can stop workspaces customising pools or using the starter pool.
- Customized workspace pools
A Spark Compute option at capacity level, enabled unless turned off, which allows Admins of the workspaces on that capacity to define their own custom Spark pools.
- Default pool
Unless an environment says otherwise, this is where a workspace's notebooks and Spark job definitions get their Spark compute. It starts out as the starter pool, and an Admin can point it at a custom pool through Workspace settings.