Prompt Shields check that defends against indirect (cross-prompt) injection by scanning third-party material, such as emails, grounding data, OCR text or tool output, for planted instructions.
Also called document attack, indirect attack, XPIA.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Prompt Shields for documents in context, with comparison tables and the common traps.
Terms in this definition
- Prompt Shields
Guardrail available in Azure AI Content Safety and Foundry for catching adversarial input. It covers jailbreaks in user prompts as well as document attacks buried in third-party content like tool responses.
- Grounding
Putting relevant and up-to-date facts into the prompt so that the model bases its answer on them, which cuts fabrication and makes citations possible. It is central to RAG.
- OCR
Optical character recognition. It reads handwritten or printed text from documents and images, making it available for analysis, something Azure Language cannot do with scanned images by itself.
- Chat message roles
Labels on chat messages: instructions go under system, the person's input under user, the model's previous answers under assistant, and results returned by a called tool under tool (or function).
Related terms
- Document attack
Instructions concealed in outside content such as emails, web pages, documents or text read from images, aiming to take over the model's session. Prompt Shields for documents detects them.
- Indirect attack evaluator
Risk and safety evaluator, also called XPIA, that checks whether an answer was hijacked by jailbreak instructions planted in retrieved documents or other context.
- XPIA
Short for cross-prompt injection attack: sneaking instructions into something an AI processes, like an email or file, so it behaves badly. Copilot screens incoming prompts with jailbreak and XPIA classifiers and stops risky ones early.