The pull-model crawler in Azure AI Search that takes data from a single data source, can apply a skillset and loads the output into a single index. It runs on demand or on a schedule, at most every five minutes.
Also called search indexer.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Indexer in context, with comparison tables and the common traps.
Terms in this definition
- Azure AI Search
Managed Azure service that builds indexes over your content and answers keyword, vector, hybrid and semantically ranked queries, optionally with AI enrichment. Retrieval-augmented generation, Foundry agents and knowledge mining all use it to fetch relevant content.
- Search data source
Connection object in Azure AI Search pointing an indexer at its content, which must be a supported source: Blob, ADLS Gen2, Table Storage, Cosmos DB, Azure SQL, SQL Managed Instance or SQL Server on Azure VMs. One indexer, one data source.
- APPLY
Evaluates a table-valued expression for every row on its left, inside
FROM. Think ofOUTER APPLYas a left outer join andCROSS APPLYas an inner join. - AI enrichment
Indexing-time process in Azure AI Search where skills such as OCR, key phrase extraction, entity recognition and image analysis turn raw content into fields that can be searched.
- Index
Speeds up queries that filter on certain columns by keeping those columns sorted, with pointers back to each row; the price is more storage and slower writes.
- Agents (classic) API
First-generation Foundry Agent Service API, based on threads, messages and runs. It is deprecated, replaced by conversations and responses, and retires on 31 March 2027.
Related terms
- Billable Foundry resource (skillset)
Key or identity of a Foundry resource linked to a skillset, letting built-in Vision, Language and Translator skills process more than the free allowance of 20 documents per indexer per day; a single multi-service resource is enough for all of them.
- Document cracking
The opening step of an indexer: source files or records are opened, and their text, metadata and, if requested, images are pulled out.
- Field mappings
Indexer configuration passing a source field unchanged into an index field with another name or type. It is applied once documents are cracked and before any skillset runs.
- imageAction
A parameter in indexer configuration. Setting it to
generateNormalizedImages, withdataToExtractset tocontentAndMetadata, pulls images out and normalises them into /document/normalized_images ready for OCR and image analysis. - Indexer schedule
A repeating interval for running an indexer, five minutes at the shortest, which allows a lengthy indexing job to continue from where it left off when change detection is turned on. It does not make indexing run in parallel.
- Integrated vectorization
Azure AI Search capability that lets an indexer split documents into chunks using the Text Split skill and produce embeddings via a skill like Azure OpenAI Embedding; a corresponding vectorizer on the index turns search queries into vectors.
- OCR skill
Azure AI Search built-in skill that relies on Azure Vision Read to pull printed and handwritten text out of
/document/normalized_images/*. The indexer must haveimageActionconfigured. - Output field mappings
Skills create nodes in the enriched document; this indexer setting routes them into index fields. It runs after the skillset, and enriched content is only indexed when mapped here.