Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
Foundry’s built-in evaluators each measure one property and each needs specific inputs. Choosing the right one is mostly a matter of matching what you want to measure with the data you actually have.
| Evaluator | Measures | Needs |
|---|---|---|
| Fluency | Grammar, vocabulary, sentence complexity, readability | Response only |
Coherence (CoherenceEvaluator) |
Logical flow and organisation | Query + response |
Relevance (RelevanceEvaluator) |
Answers the query | Query + response |
| Groundedness | Supported by retrieved context | Context |
Similarity (SimilarityEvaluator) |
Match to expected answer | Ground truth |
QAEvaluator |
Composite incl. groundedness and similarity | Context + ground truth |
Hate and unfairness (with violence, sexual, self-harm = ContentSafetyEvaluator) |
Severity of harmful content | Query + response |
| Protected material | Copyrighted text (lyrics, articles) | Query + response |
| Indirect attack | Jailbreaks injected via retrieved content | Query + response |
When a quality metric must be computed with no external knowledge (no retrieved context and no ground truth), Coherence is the usual choice: it judges the logical flow of the response against the query alone. Relevance also needs only the query and response, and Fluency only the response, so read the requirement carefully: Coherence for logical organisation, Relevance for whether the query was actually answered.
This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
Experiments, runs, parameters, metrics, artifacts and autologging.
Search spaces, sampling methods and early termination policies.
Blue-green deployments, traffic splitting, mirroring and instant rollback.
Where prompts are processed for each deployment type, and which one meets data residency needs.
When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.