How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.
From Ultra Transcenders AI-901 by Tony Rough (publishing soon)
Generative AI creates new content, so it brings its own risks: harmful, offensive or misleading output. Microsoft organises the protections into layers, each owned by a different part of the solution.
Think of the layers as a stack, from the model at the bottom to the user interface at the top. A well-designed solution uses all of them rather than relying on one. Figure 1.1 shows the four layers as a stack.
| Mitigation layer | What it is |
|---|---|
| Model | Choosing (and fine-tuning) the model |
| Safety system | Platform guardrails (content filters) that classify and block harmful prompts and completions; abuse monitoring |
| Metaprompt and grounding | System message plus grounding data in the prompt to steer the model away from harmful output |
| User experience | UI design, disclosure and constraints for users |
Common trap: The safety system layer is sometimes described as “system inputs and context”. That describes the metaprompt and grounding layer: system messages plus grounding. The safety system layer is the platform’s filtering and monitoring.
Guardrails are the main safety system control you configure yourself. They inspect what goes into and comes out of a model and act on harmful content.
Common trap: A system prompt that tells the model to avoid harmful content steers it but guarantees nothing, and fine-tuning or changing the model is not a safety control. For guaranteed blocking, configure the guardrail thresholds.
Abuse monitoring works at a different level from guardrails: it looks at behaviour over time rather than at a single response.
The same detection technology is also available as a standalone service you can call from any application.
Common trap: PII detection is not a content moderation tool. To find harmful content (hate, sexual, violence, self-harm), use Azure AI Content Safety or guardrails.
It’s a common misconception that responsible AI requires training your own model. Responsible AI with pretrained models is achieved through guardrails, grounding, evaluation, human oversight and transparency to users.
This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-901 glossary · All AI-901 study notes
Fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability, and how to tell them apart in a scenario.
What happens between a prompt and a response, and how inference differs from training.
Which model setting controls randomness, which controls length and cost, and which ones are not set at deployment.
How to recognise each AI workload from a scenario.
Agents as model plus instructions, knowledge and tools, and the auto, required and none tool_choice values.
Speech to text, text to speech, translation, batch transcription and speaker recognition compared.
How analyzers turn documents, images, audio and video into structured JSON, and how to call them from code.