How the Functions hosting plans differ in scaling, networking and cold start, and which to choose.
From Ultra Transcenders AI-200 by Tony Rough (publishing soon)
The hosting plan decides scaling, instance resources, networking, container support and cost, and it can’t be changed in place for Flex Consumption.
| Plan | Status | OS / containers | Scale-out | Max instances |
|---|---|---|---|---|
| Flex Consumption | GA, recommended serverless plan | Linux, code only | Event-driven, per-function | 1,000 per function group |
Premium (Elastic Premium, EP SKUs) |
GA | Linux code or container; Windows code | Event-driven with prewarmed workers | Windows 100; Linux 20-100 by region |
| Dedicated (App Service plan) | GA | Linux code or container; Windows code | Manual or autoscale | 10-30 (100 in an ASE) |
| Container Apps | GA | Linux containers only | Event-driven (replicas) | 300-1,000 (default 10) |
| Consumption | Legacy | Windows code; Linux retired | Event-driven | Windows 200; Linux 100 |
The Consumption plan is legacy: new serverless apps should use Flex Consumption. Linux Consumption gets no new features or language versions and retires on 30 September 2028; apps still on the end-of-life v3 runtime on Linux Consumption stopped running after 30 September 2026.
| Plan | Default functionTimeout |
Maximum |
|---|---|---|
| Flex Consumption | 30 min | Unbounded (60-minute grace on scale-in) |
| Premium | 30 min | Unbounded (60-minute grace on scale-in) |
| Dedicated | 30 min | Unbounded (requires Always On) |
| Container Apps | 30 min | Unbounded (with minimum replicas of one or more) |
| Consumption | 5 min | 10 min |
functionTimeout is set in host.json in [d.]hh:mm:ss format (-1 means unbounded). The 230-second HTTP response limit applies on every plan.
az functionapp create --resource-group <RG> --name <APP> \
--storage-account <STORAGE> --flexconsumption-location <REGION> \
--runtime python --runtime-version 3.12 \
--instance-memory 2048 --maximum-instance-count 100P (such as P1V2) is a Dedicated plan and doesn’t scale elastically.The Dedicated plan runs on App Service plan instances at App Service rates, suiting predictable billing, existing underused plans or very large compute; Always On must be enabled for unbounded timeouts and to keep non-HTTP triggers running. Container Apps hosting runs a containerised function app with the Functions programming model alongside other microservices, with GPU options; it uses revisions rather than slots (see Chapter 3, Azure Container Apps: environments, revisions and secrets).
Consumption and Flex apps can scale to zero, so the first request after idle pays a start-up delay. Flex reduces this and adds always ready instances; Premium avoids it with always ready plus prewarmed instances; Dedicated runs continuously; Container Apps avoids it when minimum replicas is one or more.
Common trap: Choosing the Consumption plan for a new Linux or Python serverless app - Consumption is legacy and Linux Consumption is retiring; Flex Consumption is the recommended serverless plan and adds virtual network support and always ready instances.
Common trap: Creating a
P1V2plan to get Premium Functions features - SKUs starting withPare Dedicated App Service plans with no elastic scale; Elastic Premium plans useEP1-EP3.
This note is one section of Ultra Transcenders AI-200: Developing AI Cloud Solutions on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · AI-200 terms in the glossary · All AI-200 study notes
The five Cosmos DB consistency levels, their RU and latency trade-offs, and when to choose each.
How to tell commands, discrete events and telemetry streams apart and pick the right Azure messaging service.
How to define KEDA scalers for queues, topics and other event sources in Container Apps.
What creates a new revision, and how single and multiple revision modes change deployments.
Exact search versus approximate IVFFlat, HNSW and DiskANN indexes, and how to tune each for recall and latency.
The main Redis caching patterns, how to expire and invalidate entries, and the trade-offs of each.
Control plane versus data plane, Azure RBAC versus vault access policies, and the roles apps need.