Vector fields and profiles, HNSW vs exhaustive KNN, hybrid queries with RRF, and the semantic ranker.
From Ultra Transcenders AI-103 by Tony Rough (publishing soon)
Vector search finds content by meaning rather than exact words, and hybrid search combines it with keyword search. Both depend on how vector fields, algorithms and vectorizers are configured in the index.
Collection(Edm.Single) with dimensions equal to the embedding model’s output (1,536 for text-embedding-ada-002), plus vectorSearchProfile. It must be searchable. It can’t be filterable, facetable or sortable, and can’t have analyzers or synonym maps.vectorSearch section: algorithms (hnsw or exhaustiveKnn), vectorizers, compressions (scalar int8 or binary quantisation) and profiles. A profile bundles an algorithm, a vectorizer and compression, and each field uses one profile.m 4 (range 4–10), efConstruction 400 and efSearch 500 (each 100–1,000). Use metric cosine with Azure OpenAI embeddings.| HNSW | Exhaustive KNN | |
|---|---|---|
| Result | Approximate nearest neighbours | Exact, brute force over all vectors |
| Memory | Graph held in memory, uses vector index quota | No vector index quota |
| Suits | Most workloads, large data | Small or medium data, precision first, ground truth for recall testing |
| Switch | "exhaustive": true per query on an HNSW field |
Can’t run HNSW queries |
Common trap: indexing a field with
exhaustiveKnnand planning to query it with HNSW later - only HNSW fields can switch per query, and only to exhaustive.
vectorQueries takes kind: "vector" (a raw array) or kind: "text".kind: "text" needs a vectorizer on the field that uses the same embedding model as indexing; this is integrated vectorisation at query time.fields (up to 10), k, and weight (default 1.0).search text plus vectorQueries in one request. The two run in parallel and merge with Reciprocal Rank Fusion (RRF). Figure 12.2 follows a hybrid query from the request to the model.@search.score values are small (about 0.03 can be a strong match).hybridSearch.maxTextRecallSize (preview) caps the BM25 side, with a default of 1,000 and a maximum of 10,000.
This note is one section of Ultra Transcenders AI-103: Developing AI Apps and Agents on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-103 glossary · All AI-103 study notes
Where guardrails check an agent run, what Prompt Shields catch, and when to block or annotate.
File search, Azure AI Search, Bing grounding, function, OpenAPI, MCP and code interpreter tools compared.
How to keep humans in the loop and limit what an agent's tools can do.
System messages, few-shot examples, chain of thought and grounding, and which fix suits which prompt problem.
The indexer pipeline stages, built-in and custom skills, and knowledge store projections.
Prebuilt and custom analyzers, field extraction methods and Markdown output for RAG.
What Sora 2 can generate, its parameters and limits, and how the asynchronous jobs work.