Brings offline evaluation of Foundry agents into CI/CD on GitHub. A test dataset is sent to the agents, catalogue evaluators grade what comes back, versions are compared statistically, and the result can hold back a release.
Also called ai-agent-evals.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Microsoft Foundry AI agent evaluation GitHub Action in context, with comparison tables and the common traps.
Terms in this definition
- CI/CD
Pipelines, for example in Azure Pipelines or GitHub Actions, that build, test and deploy code automatically; short for continuous integration and continuous delivery.
- Semantic model
Sometimes called an OLAP model, this Power BI and Fabric layer defines the measures, hierarchies, tables and relationships that reports and dashboards query; star schema design is the usual pattern.
- Built-in evaluators
Ready-made scorers in Foundry covering safety, quality, similarity, RAG and agent behaviour, with OpenAI graders and custom options alongside; they report how outputs perform but never alter settings or retrain anything.
See Microsoft Foundry AI agent evaluation GitHub Action in the full glossary