FREE STUDY NOTES · AI-901

Responsible generative AI: guardrails and the four mitigation layers

How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.

From Ultra Transcenders AI-901 by Tony Rough (publishing soon)

Generative AI creates new content, so it brings its own risks: harmful, offensive or misleading output. Microsoft organises the protections into layers, each owned by a different part of the solution.

The four mitigation layers

Think of the layers as a stack, from the model at the bottom to the user interface at the top. A well-designed solution uses all of them rather than relying on one. Figure 1.1 shows the four layers as a stack.

Four stacked layers. From the bottom up they are the model (choosing and fine-tuning it), the safety system (platform guardrails that classify and block harmful prompts and completions, plus abuse monitoring), metaprompt and grounding (a system message and grounding data that steer the model), and the user experience (UI design, disclosure and constraints for users). A note says a well-designed solution uses all four layers.
Figure 1.1: The four harm mitigation layers, from the model up to the user experience
Mitigation layer What it is
Model Choosing (and fine-tuning) the model
Safety system Platform guardrails (content filters) that classify and block harmful prompts and completions; abuse monitoring
Metaprompt and grounding System message plus grounding data in the prompt to steer the model away from harmful output
User experience UI design, disclosure and constraints for users

Common trap: The safety system layer is sometimes described as “system inputs and context”. That describes the metaprompt and grounding layer: system messages plus grounding. The safety system layer is the platform’s filtering and monitoring.

Guardrails (content filtering)

Guardrails are the main safety system control you configure yourself. They inspect what goes into and comes out of a model and act on harmful content.

Common trap: A system prompt that tells the model to avoid harmful content steers it but guarantees nothing, and fine-tuning or changing the model is not a safety control. For guaranteed blocking, configure the guardrail thresholds.

Abuse monitoring

Abuse monitoring works at a different level from guardrails: it looks at behaviour over time rather than at a single response.

Azure AI Content Safety

The same detection technology is also available as a standalone service you can call from any application.

Common trap: PII detection is not a content moderation tool. To find harmful content (hate, sexual, violence, self-harm), use Azure AI Content Safety or guardrails.

You don’t need your own model to be responsible

It’s a common misconception that responsible AI requires training your own model. Responsible AI with pretrained models is achieved through guardrails, grounding, evaluation, human oversight and transparency to users.

Get the whole book

This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · Free AI-901 glossary · All AI-901 study notes

More AI-901 study notes