A way of attacking generative AI by slipping malicious instructions into a prompt or into content the model processes, such as a file, message or web page, so that it ignores its safeguards. Azure AI Content Safety offers Prompt Shields to catch these attempts.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Prompt injection in context, with comparison tables and the common traps.
Terms in this definition
- Generative AI
AI that produces new content such as code, text, audio or images in response to a prompt. Predicting a class or a numeric value is predictive machine learning rather than generative AI.
- Azure AI Content Safety
Foundry Tool that scores text and images for hate, sexual, violent and self-harm content, returning a severity level for each category; Foundry guardrails are built on the same technology.
- Prompt Shields
Guardrail available in Azure AI Content Safety and Foundry for catching adversarial input. It covers jailbreaks in user prompts as well as document attacks buried in third-party content like tool responses.
Related terms
- Content Safety text and image moderation
Scores a piece of text or an image for severity across four harms (hate, sexual, violence, self-harm) in Azure AI Content Safety; prompt injection is outside its scope.