How analyzers turn documents, images, audio and video into structured JSON, and how to call them from code.
From Ultra Transcenders AI-901 by Tony Rough (publishing soon)
Content Understanding combines OCR and layout analysis with generative models so that you describe the fields you want and get them back as structured output. It works across documents, images, audio and video. Figure 9.1 shows the path from content to structured result.
An analyzer is the configuration that processes one kind of content.
| Analyzer | Content | Notes |
|---|---|---|
| Document (prebuilt-document, prebuilt-invoice, prebuilt-read, prebuilt-layout) | PDFs (including scanned), Office files, forms | OCR and layout, then fields |
| Image | Still images | Not PDFs |
| Audio (prebuilt-audio, prebuilt-audioSearch, prebuilt-callCenter) | Voicemails, calls | Transcription to WebVTT with diarization, speaker roles, optional word timestamps, then fields |
| Video | Video files | Keyframes and scenes |
fieldSchema defines what is extracted. For each field you give a name, a type (string, number, date, object, array and so on), a description and a method (extract, classify or generate).begin_analyze() returns an LROPoller at once, and poller.result() polls until the job completes and returns the AnalysisResult. Results are in the JSON body, not in headers.status() only reports the job’s state; wait() blocks but returns nothing; get_results doesn’t exist.This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-901 glossary · All AI-901 study notes
Fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability, and how to tell them apart in a scenario.
How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.
What happens between a prompt and a response, and how inference differs from training.
Which model setting controls randomness, which controls length and cost, and which ones are not set at deployment.
How to recognise each AI workload from a scenario.
Agents as model plus instructions, knowledge and tools, and the auto, required and none tool_choice values.
Speech to text, text to speech, translation, batch transcription and speaker recognition compared.