A Spark SQL GROUP BY extension that returns totals for each possible combination of the named columns, which is 2^n grouping sets once the overall total is counted. ROLLUP, by contrast, returns only the hierarchy-based subset.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains CUBE in context, with comparison tables and the common traps.
Terms in this definition
- Serverless
Compute tier for single Azure SQL databases that scales automatically, pauses when idle and charges by the second. It is offered in General Purpose and Hyperscale, not Business Critical, and reserved capacity does not apply.
- GROUP BY
Collapses rows sharing the same values of chosen expressions into one row per bucket, with aggregate functions evaluated for each bucket. Writing
GROUP BY ALLpicks every expression in the select list that isn't an aggregate. - GROUPING SETS
Lets one
GROUP BYquery return results for several different groupings at once, giving the same output as stitching individual aggregations together withUNION ALL.CUBEandROLLUPare convenient shorthands built on it. - ROLLUP
An option for GROUP BY, available in both Spark SQL and T-SQL, that produces subtotals at each level of a hierarchy plus an overall total; for example ROLLUP(a, b) groups by (a, b), then by (a), then by nothing. CUBE goes further and covers every combination of the columns.