Guardrail available in Azure AI Content Safety and Foundry for catching adversarial input. It covers jailbreaks in user prompts as well as document attacks buried in third-party content like tool responses.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Prompt Shields in context, with comparison tables and the common traps.
Terms in this definition
- Azure AI Content Safety
Foundry Tool that scores text and images for hate, sexual, violent and self-harm content, returning a severity level for each category; Foundry guardrails are built on the same technology.
- Chat message roles
Labels on chat messages: instructions go under system, the person's input under user, the model's previous answers under assistant, and results returned by a called tool under tool (or function).
- LIKE
Compares strings with a pattern that can contain the % and _ wildcards. Because it only understands character patterns, searching big volumes of text this way is much slower than using full-text search.
Related terms
- API Management llm-content-safety policy
Policy for the AI gateway in API Management: it passes prompts or completions to Azure AI Content Safety and rejects the call with 403 if a blocklist matches, Prompt Shields detect an attack, or a harm category passes its threshold, where 0 is strictest and 7 most lenient.
- Microsoft Defender for AI Services
Plan within Defender for Cloud that uses threat intelligence and Prompt Shields to alert in real time when generative AI applications come under attack.
- Prompt injection
A way of attacking generative AI by slipping malicious instructions into a prompt or into content the model processes, such as a file, message or web page, so that it ignores its safeguards. Azure AI Content Safety offers Prompt Shields to catch these attempts.
- Prompt Shields for documents
Prompt Shields check that defends against indirect (cross-prompt) injection by scanning third-party material, such as emails, grounding data, OCR text or tool output, for planted instructions.
- Prompt Shields for user prompts
The Prompt Shields check for direct attacks within the user's input, such as attempts to change rules, role-play personas, mocked-up conversations and encoding tricks. Documents and images are not scanned.
- Spotlighting
Encodes document content in base64 so the model gives it less trust than system or user prompts. This Prompt Shields option is in preview, off unless enabled, limited to Chat Completions and costs extra tokens.