Blue-green deployments, traffic splitting, mirroring and instant rollback.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
Updating a production model should never mean downtime or a risky cut-over. Managed online endpoints support a blue-green pattern in which the old and new models run side by side behind one address.
az ml online-endpoint update --traffic "blue=100 green=0".
Common trap: A separate staging endpoint, a new registry version, or deleting and redeploying - a staging endpoint gets no production traffic and forces clients to change URLs, a registry version doesn’t route traffic, and deleting and redeploying is slow. Shift traffic between deployments on the same endpoint instead.
This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
Experiments, runs, parameters, metrics, artifacts and autologging.
Search spaces, sampling methods and early termination policies.
Where prompts are processed for each deployment type, and which one meets data residency needs.
When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.
Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.