Deployment option that reserves model capacity for you in provisioned throughput units (PTUs) to give predictable latency and throughput. The capacity is charged by the hour regardless of usage.
Also called PTU, provisioned throughput units.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Provisioned throughput in context, with comparison tables and the common traps.
Terms in this definition
- Model deployment
Inside a resource, each deployment is a named copy of a Foundry model with its own TPM quota and deployment type. Requests identify the model by this deployment name.
- Capacity
A reserved block of compute for a tenant whose size is fixed by its SKU; Fabric F SKUs express it in capacity units (CUs). Workspaces assigned to it unlock features like Copilot, and from F64 upwards people with free licences can view Power BI content.
- total_tokens
The usage field that sums prompt_tokens and completion_tokens, for example 37 + 86 = 123. Billing applies to both kinds of token.
Related terms
- Global Provisioned
Deployment type that reserves PTU capacity and allows inference to run in any Azure region. Processing is therefore not kept within the resource's geography, as it would be with Standard.
- Serverless API deployment
You call the model as an API while Microsoft hosts it on shared infrastructure, paying per token or by PTU. This Foundry option, with types like Global Standard, Standard and Provisioned, covers Foundry Models sold by Azure, Azure OpenAI included.