When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
Standard deployments bill per token and share capacity; provisioned deployments reserve capacity for you. This section covers when to reserve, how quota and capacity differ, sizing, billing and what happens when the reservation is full.
| Deployment type | sku-name |
Minimum / increment (for example gpt-4.1, gpt-5) | Mini models (gpt-4.1-mini, gpt-5-mini, o4-mini) |
|---|---|---|---|
| Global Provisioned | GlobalProvisionedManaged |
15 / 5 | 15 / 5 |
| Data Zone Provisioned | DataZoneProvisionedManaged |
15 / 5 | 15 / 5 |
| Regional Provisioned | ProvisionedManaged |
50 / 50 | 25 / 25 |
retry-after / retry-after-ms headers.max_tokens close to the real output length, because the service reserves that amount up front.Common trap: treating a 429 from a provisioned deployment as a service fault - it is the designed signal that the PTUs are fully used; retry after the header’s wait time or redirect the request with spillover.
Spillover sends overflow traffic to a standard deployment in the same resource, so bursts above the reservation are served per token rather than rejected.
spilloverDeploymentName, or per request with the x-ms-spillover-deployment header. If both are set, the deployment property wins.IsSpillover = True on the standard deployment.This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
Experiments, runs, parameters, metrics, artifacts and autologging.
Search spaces, sampling methods and early termination policies.
Blue-green deployments, traffic splitting, mirroring and instant rollback.
Where prompts are processed for each deployment type, and which one meets data residency needs.
Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.