Foundry Tool that scores text and images for hate, sexual, violent and self-harm content, returning a severity level for each category; Foundry guardrails are built on the same technology.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Azure AI Content Safety in context, with comparison tables and the common traps.
Terms in this definition
- Chat message roles
Labels on chat messages: instructions go under system, the person's input under user, the model's previous answers under assistant, and results returned by a called tool under tool (or function).
Related terms
- Adult content detection
Image Analysis 3.2 capability, needing no training, that flags images through isAdultContent, isRacyContent and isGoryContent with scores between 0 and 1. Newer moderation is offered by Azure AI Content Safety.
- API Management llm-content-safety policy
Policy for the AI gateway in API Management: it passes prompts or completions to Azure AI Content Safety and rejects the call with 403 if a blocklist matches, Prompt Shields detect an attack, or a harm category passes its threshold, where 0 is strictest and 7 most lenient.
- Azure Content Moderator
Older moderation service, reached through the /contentmoderator path, that flagged content as a yes/no result and supported custom term lists. It is deprecated, retiring on 15 March 2027, and Azure AI Content Safety replaces it.
- Content Moderator
Older moderation service that Azure AI Content Safety has replaced; deprecated in February 2024, it retires on 15 March 2027.
- Content Safety text and image moderation
Scores a piece of text or an image for severity across four harms (hate, sexual, violence, self-harm) in Azure AI Content Safety; prompt injection is outside its scope.
- Groundedness Pro evaluator
Preview evaluator for RAG that calls Azure AI Content Safety, so no model deployment is required, and gives a strict yes or no on whether an answer agrees with the context retrieved.
- Prompt injection
A way of attacking generative AI by slipping malicious instructions into a prompt or into content the model processes, such as a file, message or web page, so that it ignores its safeguards. Azure AI Content Safety offers Prompt Shields to catch these attempts.
- Prompt Shields
Guardrail available in Azure AI Content Safety and Foundry for catching adversarial input. It covers jailbreaks in user prompts as well as document attacks buried in third-party content like tool responses.