Summary queries over a large DirectQuery table can be answered from memory instead of the source thanks to this semantic model feature: it keeps a concealed summary table cached and sends qualifying queries there, while anything needing fine detail still goes back to the source.
Also called user-defined aggregations.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Aggregations in context, with comparison tables and the common traps.
Terms in this definition
- OVER
Gives a T-SQL window function its window: PARTITION BY, ORDER BY and, if wanted, a ROWS or RANGE frame. Rankings and running totals can then be worked out while every row is kept.
- DirectQuery
Power BI connection mode in which reports fetch data from the source in real time; Direct Lake outperforms it.
- Event
Table in Log Analytics where entries from Windows event logs are kept.
- Semantic model
Sometimes called an OLAP model, this Power BI and Fabric layer defines the measures, hierarchies, tables and relationships that reports and dashboards query; star schema design is the usual pattern.
- Aggregate table
A summarised copy of a fact table at a coarser grain or with fewer dimensions (daily store sales, for instance) that makes frequent queries faster. Power BI semantic models can achieve the same result through user-defined aggregations.
- Image detail level
An option on an
input_imageitem, either low, high or the default auto, that balances resolution against token consumption. With low, a 512x512 version of the image is processed.
Related terms
- Automatic aggregations
For DirectQuery semantic models, this feature studies the query log with machine learning and then creates and maintains in-memory summaries by itself; any aggregations you've defined manually continue to work next to them.
- Dynamic M query parameters
Connects a column in the model with an M parameter, so that when a viewer picks something in a slicer or filter, that choice is written into the DirectQuery source query. It can't be used alongside row-level security or aggregations.
- Formula engine
Plans how a DAX query will be answered inside a semantic model, then asks the storage engine, VertiPaq for Import and Direct Lake, to fetch data and do the first aggregations.
- GROUPING SETS
Lets one
GROUP BYquery return results for several different groupings at once, giving the same output as stitching individual aggregations together withUNION ALL.CUBEandROLLUPare convenient shorthands built on it. - Shuffle
When Spark has to move rows between executors, as joins, sorts and aggregations require. Per stage, the Spark UI lists Shuffle Read and Shuffle Write; heavy shuffling frequently causes spill.
- Spill
Spark writing data to disk because shuffles, joins, sorts or aggregations have exhausted execution memory, which is costly. It is reported in the stage details of the Spark UI.
- Stateful operation
Any Structured Streaming operation that must hold intermediate data between batches (stream-stream joins, deduplication or aggregations over windows, for instance). That data sits in a state store kept with the checkpoint, and TTL, watermarks and time-range conditions keep it from growing endlessly.
- SUMMARIZE
Groups a table in DAX, giving a row for every combination of the columns chosen to group by, and can add named columns that hold aggregations.