Risk and safety evaluator, also called XPIA, that checks whether an answer was hijacked by jailbreak instructions planted in retrieved documents or other context.
Also called XPIA (cross-domain prompt injected attack).
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Indirect attack evaluator in context, with comparison tables and the common traps.
Terms in this definition
- Risk
As ISO 31000 puts it, how uncertainty affects objectives; that effect can be good or bad.
- Prompt Shields for documents
Prompt Shields check that defends against indirect (cross-prompt) injection by scanning third-party material, such as emails, grounding data, OCR text or tool output, for planted instructions.
- Prompt Shields for user prompts
The Prompt Shields check for direct attacks within the user's input, such as attempts to change rules, role-play personas, mocked-up conversations and encoding tricks. Documents and images are not scanned.