Prebuilt and custom analyzers, field extraction methods and Markdown output for RAG.
From Ultra Transcenders AI-103 by Tony Rough (publishing soon)
Content Understanding processes documents, images, audio and video through analyzers that combine classic extraction with generative field extraction. Configure the analyzer first, then let the agent reason over its output.
| Analyzer | Does |
|---|---|
prebuilt-read |
Text and barcodes, no layout, no model needed |
prebuilt-layout |
Words, paragraphs, figures, tables, sections, barcodes (QRCode, MicroQRCode); no LLM deployment |
prebuilt-invoice and other prebuilt domain analyzers |
Fields across varied layouts, no training |
prebuilt-documentSearch |
RAG analyzer using generative models |
prebuilt-documentFieldSchema |
Proposes a field schema |
Custom analyzer (based on prebuilt-document or a copied template) |
Your schema, fields and validation (for example, checking against contract terms) |
estimateFieldSourceAndConfidence: true on the analyzer, or estimateSourceAndConfidence per field.enableSegment splits documents for classification.search.score measure other things, not extraction confidence.PUT {endpoint}/contentunderstanding/analyzers/{analyzerId}?api-version=..., which returns 201 plus Operation-Location.POST ...:analyze runs an analyzer, GET reads one and PATCH updates one.Operation-Location until the status is Succeeded. The GA API (2025-11-01) keeps this 202 plus Operation-Location pattern.Common trap: using POST to create an analyzer - creation is PUT to the analyzer’s URL; POST with
:analyzeruns it.
This note is one section of Ultra Transcenders AI-103: Developing AI Apps and Agents on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-103 glossary · All AI-103 study notes
Where guardrails check an agent run, what Prompt Shields catch, and when to block or annotate.
File search, Azure AI Search, Bing grounding, function, OpenAPI, MCP and code interpreter tools compared.
How to keep humans in the loop and limit what an agent's tools can do.
System messages, few-shot examples, chain of thought and grounding, and which fix suits which prompt problem.
The indexer pipeline stages, built-in and custom skills, and knowledge store projections.
Vector fields and profiles, HNSW vs exhaustive KNN, hybrid queries with RRF, and the semantic ranker.
What Sora 2 can generate, its parameters and limits, and how the asynchronous jobs work.