Quality and safety scoring of models and agents in Microsoft Foundry, using evaluators that are either built in or custom. Evaluations measure risk but do not block anything at runtime.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Foundry evaluations in context, with comparison tables and the common traps.
Terms in this definition
- Microsoft Foundry
Formerly called Azure AI Foundry, the Azure platform where AI models and agents are built, evaluated and run within one resource, with evaluations, guardrails and AI red teaming.
- Built-in evaluators
Ready-made scorers in Foundry covering safety, quality, similarity, RAG and agent behaviour, with OpenAI graders and custom options alongside; they report how outputs perform but never alter settings or retrain anything.
- Measure
A figure like total revenue, worked out by aggregating fact table data wherever dimensions intersect within a semantic model.
- Risk
As ISO 31000 puts it, how uncertainty affects objectives; that effect can be good or bad.
Related terms
- OpenAI graders
Evaluators of four kinds, string check, text similarity, score model and label model, available both in Foundry evaluations and in the GitHub Action for evaluation.