The indexer pipeline stages, built-in and custom skills, and knowledge store projections.
From Ultra Transcenders AI-103 by Tony Rough (publishing soon)
A skillset is the chain of AI skills an indexer runs on each document. Skills chunk and vectorise text, read images, translate and extract entities, and custom skills call your own code.
imageAction: generateNormalizedImages (with dataToExtract: contentAndMetadata). The images land in /document/normalized_images, and the OCR skill uses /document/normalized_images/*.defaultFromLanguageCode to auto-detect the source language.Microsoft.Skills.Custom.WebApiSkill with a uri, such as an Azure Function) can wrap Document Intelligence’s prebuilt invoice model to extract invoice fields.<context>/<output>, and a context ending /* runs once per item.Common trap: attaching an Azure Vision (Computer Vision) resource to bill the skillset - the Language skills aren’t covered; attach a multi-service Foundry resource.
A pipeline that processes multilingual PDFs and keeps the enriched output uses these stages:
| Stage | Component |
|---|---|
| Source | Blob Storage |
| Cracking | Vision OCR |
| Preparation | Translator |
| Destination | Blob knowledge store |
Common trap: choosing Azure Files as the enrichment destination - knowledge stores write only to Blob and Table Storage.
This note is one section of Ultra Transcenders AI-103: Developing AI Apps and Agents on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-103 glossary · All AI-103 study notes
Where guardrails check an agent run, what Prompt Shields catch, and when to block or annotate.
File search, Azure AI Search, Bing grounding, function, OpenAPI, MCP and code interpreter tools compared.
How to keep humans in the loop and limit what an agent's tools can do.
System messages, few-shot examples, chain of thought and grounding, and which fix suits which prompt problem.
Vector fields and profiles, HNSW vs exhaustive KNN, hybrid queries with RRF, and the semantic ranker.
Prebuilt and custom analyzers, field extraction methods and Markdown output for RAG.
What Sora 2 can generate, its parameters and limits, and how the asynchronous jobs work.