Where guardrails check an agent run, what Prompt Shields catch, and when to block or annotate.
From Ultra Transcenders AI-103 by Tony Rough (publishing soon)
Guardrails are the built-in, configurable safety layer that sits between users, models and tools in Microsoft Foundry. They decide what is checked, where in a run it is checked, and what happens when something is detected.
Each control catches a specific class of problem; none of them covers everything, so production solutions layer several.
| Control | Catches | Doesn’t catch |
|---|---|---|
| Prompt Shields for user prompts (was jailbreak risk detection) | Direct attacks in the user’s own input: rule changes, role-play personas, conversation mockups, encoding tricks | Instructions in documents or images; unsafe imagery |
| Prompt Shields for documents | Indirect attacks in third-party content: grounding data, OCR text, tool responses | Unsafe imagery |
| Spotlighting (preview, Prompt Shields sub-feature) | Marks document content (base64) as lower trust than system and user prompts | — |
| Text and image moderation | Hate, sexual, violence, self-harm by severity | Prompt injection |
| Protected material detection | Known copyrighted text in output (lyrics, articles, recipes, web content) and code from GitHub repos | Harm categories, injection |
| Custom blocklists | Exact listed terms | Unknown phrasing |
| PII detection | Personal data | Injection, harm |
Common trap: Choosing abuse monitoring to block hate speech in model output - abuse monitoring detects patterns of misuse for Microsoft to review; it’s content filters (guardrails) that block individual responses.
This note is one section of Ultra Transcenders AI-103: Developing AI Apps and Agents on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-103 glossary · All AI-103 study notes
File search, Azure AI Search, Bing grounding, function, OpenAPI, MCP and code interpreter tools compared.
How to keep humans in the loop and limit what an agent's tools can do.
System messages, few-shot examples, chain of thought and grounding, and which fix suits which prompt problem.
The indexer pipeline stages, built-in and custom skills, and knowledge store projections.
Vector fields and profiles, HNSW vs exhaustive KNN, hybrid queries with RRF, and the semantic ranker.
Prebuilt and custom analyzers, field extraction methods and Markdown output for RAG.
What Sora 2 can generate, its parameters and limits, and how the asynchronous jobs work.