The Prompt Shields check for direct attacks within the user's input, such as attempts to change rules, role-play personas, mocked-up conversations and encoding tricks. Documents and images are not scanned.
Also called user prompt attack, jailbreak, Jailbreak risk detection.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Prompt Shields for user prompts in context, with comparison tables and the common traps.
Terms in this definition
- Prompt Shields
Guardrail available in Azure AI Content Safety and Foundry for catching adversarial input. It covers jailbreaks in user prompts as well as document attacks buried in third-party content like tool responses.
- Chat message roles
Labels on chat messages: instructions go under system, the person's input under user, the model's previous answers under assistant, and results returned by a called tool under tool (or function).
- ENCODING
A COPY INTO option for CSV sources that states whether the files use UTF8, the default, or UTF16 character encoding.
Related terms
- Indirect attack evaluator
Risk and safety evaluator, also called XPIA, that checks whether an answer was hijacked by jailbreak instructions planted in retrieved documents or other context.
- User prompt attack
Also called a jailbreak, this is a harmful prompt typed by the user to override the system message or safety training. Prompt Shields for user prompts flags it; instructions concealed in documents are a different threat it misses.
- XPIA
Short for cross-prompt injection attack: sneaking instructions into something an AI processes, like an email or file, so it behaves badly. Copilot screens incoming prompts with jailbreak and XPIA classifiers and stops risky ones early.