Ready-made scorers in Foundry covering safety, quality, similarity, RAG and agent behaviour, with OpenAI graders and custom options alongside; they report how outputs perform but never alter settings or retrain anything.
Also called evaluators.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Built-in evaluators in context, with comparison tables and the common traps.
Terms in this definition
- RAG
Technique for answering from private or recent information: relevant content is fetched from your own data, typically a search index, and inserted into the prompt as grounding so the model can respond with citations.
- Agent
A specialised form of Microsoft Copilot set up for one particular job, pairing instructions with knowledge and skills. You can create one in Copilot Studio, SharePoint or Agent Builder, and administrators control them from the Microsoft 365 admin center.
- OpenAI graders
Evaluators of four kinds, string check, text similarity, score model and label model, available both in Foundry evaluations and in the GitHub Action for evaluation.
- Table options
Set only at creation with
OPTIONS, these key-value storage settings can't be altered or dropped afterwards; on Delta tables they also show up among the table properties.
Related terms
- Custom evaluators
When none of the built-in evaluators suit, you can write your own as an LLM-as-judge prompt or as Python code.
- Foundry evaluations
Quality and safety scoring of models and agents in Microsoft Foundry, using evaluators that are either built in or custom. Evaluations measure risk but do not block anything at runtime.
- Ground truth
The expected answer included with evaluation data, which evaluators like F1 or Response Completeness use as the benchmark when judging a model's response.
- Microsoft Foundry AI agent evaluation GitHub Action
Brings offline evaluation of Foundry agents into CI/CD on GitHub. A test dataset is sent to the agents, catalogue evaluators grade what comes back, versions are compared statistically, and the result can hold back a release.
- Risk and safety evaluators
Foundry's safety-focused evaluators, which score outputs for things like hate and unfairness, violence, sexual and self-harm content, protected material, indirect attacks, prohibited actions and leaked sensitive data. Scoring is all they do; nothing is blocked.