Where prompts are processed for each deployment type, and which one meets data residency needs.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
A deployment type decides where your prompts are processed, how much quota you get and whether requests are served interactively or in batches. Pick the type from the workload’s residency and latency requirements first, then from cost. Figure 7.1 lays the types out from the widest processing area to the tightest residency.
| Deployment type | Behaviour |
|---|---|
| Global Standard | Routes to any region with capacity; highest default quota; for workloads used across many regions |
| Standard (regional) | Processes data within the resource’s Azure geography (possibly across regions in that geography); the option when data must stay in one geography |
| Data Zone Standard | Real-time pay-per-token within a data zone (e.g. US or EU); more quota than a Standard (regional) deployment |
| Global Batch / Data Zone Batch | Asynchronous, target turnaround up to 24 hours; not interactive |
| Developer | Evaluating fine-tuned models only; no SLA, not for production |
The provisioned equivalents of Global, Data Zone and regional deployments are covered under provisioned throughput later in this chapter.
Common trap: choosing Data Zone Standard to keep a workload in one geography - a data zone spans many regions and countries; to keep processing within one Azure geography use Standard (regional).
This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
Experiments, runs, parameters, metrics, artifacts and autologging.
Search spaces, sampling methods and early termination policies.
Blue-green deployments, traffic splitting, mirroring and instant rollback.
When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.
Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.