The expected answer included with evaluation data, which evaluators like F1 or Response Completeness use as the benchmark when judging a model's response.
Also called expected response.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Ground truth in context, with comparison tables and the common traps.
Terms in this definition
- Built-in evaluators
Ready-made scorers in Foundry covering safety, quality, similarity, RAG and agent behaviour, with OpenAI graders and custom options alongside; they report how outputs perform but never alter settings or retrain anything.
- LIKE
Compares strings with a pattern that can contain the % and _ wildcards. Because it only understands character patterns, searching big volumes of text this way is much slower than using full-text search.
Related terms
- Evaluation dataset
A constant collection of test inputs, a query plus optional context and ground truth, fed to every model or prompt variant so their results can be compared fairly.
- QAEvaluator
All-in-one evaluator returning F1, similarity, fluency, coherence, relevance and groundedness scores together. You must pass in context and a ground truth for it to run.
- Response Completeness evaluator
Evaluator in preview for RAG scenarios: it asks whether the answer contains the essential facts from the ground truth, so it measures recall, complementing groundedness, which measures precision.
- Similarity evaluator
Evaluator that rates how textually close a response is to the expected answer, so ground truth must be supplied.