Deployment type billed per token whose data may be processed in any Azure region that has capacity. It offers the biggest default quota and earliest access to new models, making it the recommended first choice.
Also called GlobalStandard.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Global Standard in context, with comparison tables and the common traps.
Terms in this definition
- Model deployment
Inside a resource, each deployment is a named copy of a Foundry model with its own TPM quota and deployment type. Requests identify the model by this deployment name.
- Token
The unit of text an LLM works with, which may be a word, part of a word or punctuation. Billing, limits and context windows are all counted in these units.
- region
A provider-defined grouping in VCF Automation of Supervisors that all share one NSX Local Manager; tenants consume its compute, storage and memory via quotas set per region.
- Capacity
A reserved block of compute for a tenant whose size is fixed by its SKU; Fabric F SKUs express it in capacity units (CUs). Workspaces assigned to it unlock features like Copilot, and from F64 upwards people with free licences can view Power BI content.
- FIRST
A DAX function available only inside visual calculations. It fetches the value at the start of one axis of the visual's matrix, which makes it handy for comparing each point with the first; its opposite is LAST.
Related terms
- Foundry deployment type
Option picked at model deployment time, such as Global Standard, Standard, Data Zone, provisioned or batch. It determines both the billing model and the location where data gets processed.
- Serverless API deployment
You call the model as an API while Microsoft hosts it on shared infrastructure, paying per token or by PTU. This Foundry option, with types like Global Standard, Standard and Provisioned, covers Foundry Models sold by Azure, Azure OpenAI included.