FREE STUDY NOTES · AI-103

Guardrails in Microsoft Foundry: intervention points, Prompt Shields and actions

Where guardrails check an agent run, what Prompt Shields catch, and when to block or annotate.

From Ultra Transcenders AI-103 by Tony Rough (publishing soon)

Guardrails are the built-in, configurable safety layer that sits between users, models and tools in Microsoft Foundry. They decide what is checked, where in a run it is checked, and what happens when something is detected.

Harm filtering and severity thresholds

Intervention points and actions for agents

An agent run passes four guardrail intervention points: user input, tool call, tool response and output, each set to Annotate and block (agents don't support Annotate only). Prompt Shields for user prompts apply at user input and Prompt Shields for documents at tool response. The tool call and tool response points are in preview and apply only to supported tools.
Figure 4.1: Guardrail intervention points along an agent run

Choosing the right control

Each control catches a specific class of problem; none of them covers everything, so production solutions layer several.

Control Catches Doesn’t catch
Prompt Shields for user prompts (was jailbreak risk detection) Direct attacks in the user’s own input: rule changes, role-play personas, conversation mockups, encoding tricks Instructions in documents or images; unsafe imagery
Prompt Shields for documents Indirect attacks in third-party content: grounding data, OCR text, tool responses Unsafe imagery
Spotlighting (preview, Prompt Shields sub-feature) Marks document content (base64) as lower trust than system and user prompts —
Text and image moderation Hate, sexual, violence, self-harm by severity Prompt injection
Protected material detection Known copyrighted text in output (lyrics, articles, recipes, web content) and code from GitHub repos Harm categories, injection
Custom blocklists Exact listed terms Unknown phrasing
PII detection Personal data Injection, harm

Common trap: Choosing abuse monitoring to block hate speech in model output - abuse monitoring detects patterns of misuse for Microsoft to review; it’s content filters (guardrails) that block individual responses.

Get the whole book

This note is one section of Ultra Transcenders AI-103: Developing AI Apps and Agents on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · Free AI-103 glossary · All AI-103 study notes

More AI-103 study notes