FREE STUDY NOTES · AI-300

Provisioned throughput units (PTU) for Foundry models

When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.

From Ultra Transcenders AI-300 by Tony Rough (publishing soon)

Standard deployments bill per token and share capacity; provisioned deployments reserve capacity for you. This section covers when to reserve, how quota and capacity differ, sizing, billing and what happens when the reservation is full.

When to use provisioned throughput

Quota, capacity and sizing

Deployment type sku-name Minimum / increment (for example gpt-4.1, gpt-5) Mini models (gpt-4.1-mini, gpt-5-mini, o4-mini)
Global Provisioned GlobalProvisionedManaged 15 / 5 15 / 5
Data Zone Provisioned DataZoneProvisionedManaged 15 / 5 15 / 5
Regional Provisioned ProvisionedManaged 50 / 50 25 / 25

Billing and reservations

Utilisation and 429 responses

Common trap: treating a 429 from a provisioned deployment as a service fault - it is the designed signal that the PTUs are fully used; retry after the header’s wait time or redirect the request with spillover.

Spillover

Spillover sends overflow traffic to a standard deployment in the same resource, so bursts above the reservation are served per token rather than rejected.

Get the whole book

This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · Free AI-300 glossary · All AI-300 study notes

More AI-300 study notes