Technique for answering from private or recent information: relevant content is fetched from your own data, typically a search index, and inserted into the prompt as grounding so the model can respond with citations.
Also called retrieval augmented generation.
Read more: Microsoft Learn
In the Ultra Transcenders books
AI-901AI-300AI-103AI-200DP-800
Each book explains RAG in context, with comparison tables and the common traps.
Terms in this definition
- Index
Speeds up queries that filter on certain columns by keeping those columns sorted, with pointers back to each row; the price is more storage and slower writes.
- Grounding
Putting relevant and up-to-date facts into the prompt so that the model bases its answer on them, which cuts fabrication and makes citations possible. It is central to RAG.
Related terms
- Agentic retrieval
Query approach in Azure AI Search where an LLM turns a conversation into several planned subqueries; these execute together, are reranked semantically, and the best chunks are merged. Classic RAG, by contrast, issues just one query.
- Built-in evaluators
Ready-made scorers in Foundry covering safety, quality, similarity, RAG and agent behaviour, with OpenAI graders and custom options alongside; they report how outputs perform but never alter settings or retrain anything.
- Groundedness evaluator
An evaluator for RAG that rates, from 1 to 5, how far a response is backed by the supplied context rather than invented (the precision side). Context must be provided.
- Groundedness Pro evaluator
Preview evaluator for RAG that calls Azure AI Content Safety, so no model deployment is required, and gives a strict yes or no on whether an answer agrees with the context retrieved.
- Markdown output (Document Intelligence)
Option on the layout model that returns the extracted content as Markdown, rendering tables as HTML; handy when chunking documents for RAG.
- Relevance evaluator
Evaluator for RAG output that judges, from the query and response alone, how directly and accurately the answer deals with what the user asked.
- Response Completeness evaluator
Evaluator in preview for RAG scenarios: it asks whether the answer contains the essential facts from the ground truth, so it measures recall, complementing groundedness, which measures precision.
- Training cutoff
The point in time beyond which a model knows nothing. Grounding with RAG at request time, not retraining, is how recent events are covered.