A list of numbers, spanning many dimensions, that captures meaning; items with similar meaning end up near one another, as cosine similarity measures.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Embedding in context, with comparison tables and the common traps.
Terms in this definition
- List
Permission on Key Vault secrets allowing a caller to enumerate those in a vault, though their values are not returned.
- Cosine similarity
Scores how semantically close two embedding vectors are by the angle between them; vector search relies on it.
Related terms
- Azure OpenAI Embedding skill
Skill in an AI Search skillset that sends each chunk to an embedding model deployment (text-embedding-3-small, text-embedding-3-large or ada-002) during indexing to produce vectors; charges come from Azure OpenAI.
- Base64 data URI
Embedding the image itself in
image_urlas adata:image/jpeg;base64,...string; this is the approach when no public URL hosts the image. - CUSTOMDATA
Reads whatever an embedding application (or anything else) put in the connection string's CustomData property, and returns blank when that property is empty. This is a DAX function.
- Embeddings
Vectors of numbers encoding what text means, placing related content close together as judged by cosine similarity. Switching model or vector size requires embedding all content again.
- Integrated vectorization
Azure AI Search capability that lets an indexer split documents into chunks using the Text Split skill and produce embeddings via a skill like Azure OpenAI Embedding; a corresponding vectorizer on the index turns search queries into vectors.
- text-embedding-ada-002
An embedding model in Azure OpenAI that accepts up to 8,192 input tokens and outputs only 1,536-dimension vectors, never generated text.
- Text Split skill
A free utility skill that divides text into pages or sentences, optionally overlapping, to create chunks ready for embedding. Vectorisation happens elsewhere.