For each harm category, the content filter level (medium by default) from which content gets blocked. The strictest option is blocking at low for both input and output.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Severity threshold in context, with comparison tables and the common traps.
Terms in this definition
- Content filter
Guardrail set up in Foundry that rates prompts and completions for risks such as hate, violence, sexual content, self-harm, prompt attacks and protected material, blocking anything at or above a chosen severity.