ULTRATRANSCENDERS

AI-901 glossary

Microsoft Azure AI Fundamentals · the glossary from the book, with every entry linked to Microsoft Learn · Get the book

A

Abuse monitoring

Azure OpenAI process that detects patterns of misuse across a customer's traffic for Microsoft review; it does not filter individual responses.

Accountability principle

Responsible AI principle that named people remain answerable for AI, with human oversight, the ability to override and governance committees.

Accuracy (classification metric)

Proportion of correct predictions over all test cases, an overall model evaluation metric for classification models.

Agent instructions (system message)

The agent's system message: sets its role, goals, tone and limits and guides (non-deterministically) when it uses tools; deployment slots, embeddings or fine-tuning don't define the role.

AI agent

Application that uses a generative model with tools to interpret requests, hold a dialogue and take actions toward a goal.

AI enrichment (skillset)

Azure AI Search pipeline in which skills (OCR, entity recognition, key phrases, image analysis) transform raw content into searchable fields during indexing.

AKS (Azure Kubernetes Service)

Managed Kubernetes with full cluster and node-pool control; autoscales with HPA and the cluster autoscaler; has no built-in user sign-in.

AnalysisResult

Content Understanding result object returned by poller.result(), holding markdown content and extracted fields in the JSON body, not in headers.

Anomaly detection

Finding unusual patterns in numeric or time-series data, such as fraud or sensor spikes; the Anomaly Detector service was retired on 1 October 2026.

API (application programming interface)

Programmatic interface a client calls, such as the Responses, Images or Files API.

API key (Foundry resource key) (resource key)

Secret key issued with a Foundry Tools or Azure OpenAI resource and sent in the api-key header; the fallback when keyless Entra ID auth isn't possible.

Area of interest (smart crop)

Image Analysis feature that returns the bounding box of an image's most important region for a given aspect ratio, used for cropping and thumbnails.

Attention

Transformer mechanism that weighs how much each other token in the context influences the next prediction; it is not the embedding vector.

AudioOutputConfig

Speech SDK class that sets where synthesised audio goes: the default speaker, a file (filename="output.wav"), a custom output stream or a named device, one at a time.

AudioStreamFormat

Speech SDK class that only describes an audio stream's encoding (sample rate, bits, channels); it isn't an output destination.

Azure AI Agent Service

See Foundry Agent Service.

Azure AI Content Safety

Foundry Tool that detects hate, sexual, violence and self-harm content in text and images with severity levels; the same technology powers Foundry guardrails.

Azure AI Content Understanding

See Azure Content Understanding in Foundry Tools.

Azure AI Face (Face service)

Foundry Tool that detects, verifies (1:1) and identifies (1:many) human faces; identify and verify are Limited Access, and it handles faces only.

Azure AI Foundry

See Microsoft Foundry.

Azure AI Language

See Azure Language in Foundry Tools.

Azure AI services

See Foundry Tools.

Azure AI Speech

See Azure Speech in Foundry Tools.

Azure AI Translator

See Azure Translator in Foundry Tools.

Azure Content Understanding in Foundry Tools (Content Understanding)

Foundry Tool that uses generative AI analyzers to extract, classify and generate structured fields from documents, images, audio and video.

Azure Document Intelligence in Foundry Tools (Document Intelligence, Form Recognizer)

Foundry Tool with read, layout, prebuilt (invoice, receipt) and custom models that extract text and fields from forms and documents.

Azure Face service (Face API)

Azure Vision service for face detection, verification (one-to-one) and identification (one-to-many); Limited Access, and the only Vision option that identifies who a face belongs to.

Azure Language in Foundry Tools (Azure AI Language, Text Analytics)

Foundry Tool of NLP features (sentiment, key phrases, NER, PII, summarization, language detection) that analyse digital text; prebuilt features need no training and it doesn't generate new content.

Azure Machine Learning (AML)

Cloud service for training, deploying and managing your own ML models with MLOps; the wrong answer when a prebuilt language model does the writing.

Azure Machine Learning designer (Designer)

Drag-and-drop interface in Azure Machine Learning studio for building ML training pipelines; not where you browse or deploy Foundry models.

Azure OpenAI (Azure OpenAI in Foundry Models)

Azure access to OpenAI models (GPT, embeddings, gpt-image) through a Foundry or Azure OpenAI resource; a standalone Azure OpenAI resource doesn't expose Content Understanding.

Azure OpenAI in Foundry Models (Azure OpenAI Service)

Fully managed Azure access to OpenAI models (GPT-4.1, GPT-5, o-series, image, audio, embeddings) with built-in guardrails; you choose a deployment type instead of managing infrastructure.

Azure OpenAI Service

See Azure OpenAI in Foundry Models.

Azure OpenAI v1 API

Versionless Azure OpenAI API called with the standard OpenAI() client and base_url https://<resource-name>.openai.azure.com/openai/v1/; the host is the resource name, not the project.

Azure RBAC

Azure role-based access control: assigns built-in or custom roles to users, groups and managed identities at a scope; keyless (Microsoft Entra ID) calls to Foundry need a data-plane role, such as Foundry User (previously Azure AI User) on the project, or Cognitive Services OpenAI User for the Azure OpenAI endpoint.

Azure Speech in Foundry Tools (Azure Speech)

Foundry Tool providing speech to text, text to speech, speech translation and live voice conversations.

Azure Translator in Foundry Tools (Translator)

Foundry Tool for text and document translation between languages; it doesn't extract fields from forms.

Azure Vision in Foundry Tools (Azure AI Vision, Computer Vision)

Foundry Tool for image analysis, OCR and face detection; the Image Analysis API (v3.2 and v4.0) is deprecated and retires on 25 September 2028.

B

Base analyzer

Modality-specific parent analyzer (prebuilt-document, prebuilt-image, prebuilt-audio, prebuilt-video) that custom analyzers inherit from via baseAnalyzerId.

Base64 data URI

Inline encoding of image bytes as data:image/jpeg;base64,... in image_url, used when the image isn't at a public URL.

Batch transcription

Asynchronous speech to text REST API for large volumes of prerecorded audio in storage, with results written back later; not for live captions.

begin_analyze()

ContentUnderstandingClient method that starts asynchronous analysis and returns an LROPoller at once; call poller.result() for the AnalysisResult.

Blob storage

Object storage for unstructured data such as video and images; block blobs up to about 190.7 TiB.

Bounding box

Four pixel values (x, y, width, height) that locate a detected object in an image; returned by object detection but not by image classification.

Brand detection (logo detection)

Specialised object detection in Azure Vision Image Analysis (v3.2) that recognises commercial logos from a database of thousands of brands and returns the brand name, confidence and bounding box.

Built-in evaluators (evaluators)

Foundry evaluators that measure and report quality, RAG, safety and agent behaviour; they score outputs but don't change settings or retrain the model.

C

Captioning

Speech to text scenario that converts the audio of live or recorded video into synchronised on-screen text, with optional profanity filtering and partial results.

Chat playground

No-code Foundry page for testing a deployed chat model with a system message and parameters such as temperature, max response and past messages.

Classification

Supervised ML that predicts which category an item belongs to, such as spam or not spam; predictive, not generative.

CLU (conversational language understanding)

Azure Language custom feature that predicts the intent of user utterances and extracts entities for chatbots; retiring 31 March 2029.

Codex

Legacy OpenAI code model (code-davinci-002) retired in Azure; newer GPT models, including the GPT-5 codex variants, handle code today (codex-mini is itself deprecated and retires on 15 November 2026).

Coherence evaluator

General-purpose Foundry evaluator that scores the logical flow and consistency of a response.

Color scheme detection (colour scheme)

Image Analysis 3.2 feature that reports whether an image is black and white and its dominant foreground, background and accent colours.

Computer vision

Workload that interprets images and video, such as detecting people, faces, objects, logos or flooded areas.

Confidence score

Probability from 0 to 1 attached to each prediction or extracted field, showing how sure the model is; distinct from accuracy, which is measured over a test set.

Content extraction

First Content Understanding stage that normalises input into text and metadata: OCR and layout for documents, transcription for audio and video.

Content filters

See Foundry guardrail.

Content Understanding analyzer (analyzer)

Configuration in Content Understanding that defines how content is processed and which fields are extracted, via prebuilt analyzers or a custom field schema.

ContentUnderstandingClient

Python client for Content Understanding whose begin_analyze() starts an analysis and returns a poller.

Contract model (prebuilt-contract)

Document Intelligence prebuilt model that extracts contract fields such as Parties, Jurisdictions, Contract ID and Title.

Conversational AI

Workload of two-way dialogue where software interprets questions and replies, such as chat bots; on AI-901 framed as a generative AI chat app or agent in Foundry.

Custom extraction model

Document Intelligence model you label and train to extract your own fields; needs as few as five labelled samples and comes as custom template or custom neural.

Custom NER (custom named entity recognition)

Azure Language feature you train to recognise domain-specific entities; it finds but doesn't mask them.

Custom neural model

Deep-learning custom extraction model that generalises across varied layouts of one document type (structured and semi-structured); the recommended starting point over custom template.

Custom text classification

Azure Language feature you train to assign whole documents to classes you define, such as work or personal email.

Custom Vision (Azure AI Custom Vision)

Foundry Tool for training your own image classification and object detection models from labelled images; it extracts no document fields and is retiring (support until 25 September 2028).

Custom voice (custom neural voice, brand voice)

Limited-access text to speech feature that trains a unique synthetic brand or character voice from recorded human speech samples.

D

DALL-E (dall-e-3)

Retired OpenAI image generation model; dall-e-3 was retired in Azure on 4 March 2026 and replaced by the gpt-image series.

Diarization (speaker diarization)

Separation of a transcript by speaker, labelling who spoke when; supported by Speech and Content Understanding audio analyzers.

E

Embedding

Multi-dimensional vector of numbers representing meaning, where semantically similar items sit close together (measured by cosine similarity).

Embeddings model (text-embedding-3)

Model that converts text into vectors for semantic search, similarity and recommendations; it doesn't generate text.

Entity linking

Azure Language feature that disambiguates entities and links them to Wikipedia; it doesn't redact and retires on 1 September 2028.

F

Fabrication (hallucination)

Plausible but incorrect output that an LLM generates because it doesn't fact-check; grounding reduces but never eliminates it.

Face detection

Locating human faces in an image and returning bounding boxes; it doesn't identify or verify who the person is.

Face detection (Image Analysis)

Image Analysis 3.2 feature that returns a rectangle for each detected face; it doesn't identify or verify people (use the Face service).

Fairness assessment

Responsible AI dashboard component (built on Fairlearn) that compares model performance across sensitive groups such as gender, ethnicity and age using disparity metrics.

Fairness principle

Responsible AI principle that AI treats everyone fairly and similarly situated groups aren't affected differently; it is not the same as overall accuracy or identical output for all.

Fast transcription

Speech to text REST API that transcribes prerecorded audio files synchronously, faster than real time, with predictable latency.

Few-shot learning (few-shot prompting)

Including example input-output pairs in the prompt to show the expected format and pattern for the current request; the model's weights don't change (zero-shot uses no examples).

Field generation method (extract, classify, generate)

Per-field method in a Content Understanding schema: extract (values as written, documents only), classify (pick from categories) or generate (infer or summarise).

Field schema (fieldSchema)

Content Understanding definition of the fields to extract, giving each a name, description, type and method (extract, classify or generate).

Files API

Azure OpenAI API for uploading files; an uploaded image's file ID can be referenced in an input_image item and reused across requests.

Fine-tuning

Further training a pretrained model on a task-specific dataset to adjust its weights for style, format or task performance; not a safety control and not how to add fresh knowledge.

Fluency evaluator

General-purpose Foundry evaluator that scores the natural-language quality and readability of a response.

Form Recognizer

See Azure Document Intelligence in Foundry Tools.

Foundation model

Large model pretrained on huge unlabelled corpora (self-supervised next-token prediction) that can be used as-is or adapted for many tasks.

Foundry Agent Service

Managed Microsoft Foundry platform for building, deploying and scaling prompt, voice and hosted agents with models, tools and guardrails.

Foundry deployment type

Choice made when deploying a model (Global Standard, Standard, Data Zone, provisioned, batch) that sets where data is processed and how you pay.

Foundry evaluations

Microsoft Foundry scoring of model and agent quality and safety with built-in or custom evaluators; measures risk without blocking at runtime.

Foundry guardrail

Named collection of controls (risk, intervention point, action) in Microsoft Foundry, previously called content filters, that classifies prompts and completions for hate and fairness, sexual, violence and self-harm and blocks or annotates them.

Foundry playground (model playground)

Foundry portal environment for testing a model or deployment interactively, adjusting system prompt, temperature, top_p and max tokens, comparing models and exporting sample code; not for production traffic.

Foundry project (project)

Child of a Foundry resource that isolates a team's agents, files and evaluations while sharing the resource's deployments and connections.

Foundry Tools (Azure AI services, Azure Cognitive Services)

Family of prebuilt and customisable AI APIs (Language, Speech, Vision, Content Understanding, Document Intelligence, Content Safety and more).

Foundry Tools containers

Docker containers that run some Foundry Tools (such as Read OCR and Language) on-premises or at the edge while billing to Azure; available for prebuilt as well as custom features.

Frequency penalty (frequency_penalty)

Inference parameter (-2.0 to 2.0) that penalises tokens by how often they've appeared, mainly reducing verbatim repetition.

G

General document model (prebuilt-document)

Deprecated Document Intelligence model that extracted key-value pairs; replaced in v4.0 by the layout model with the keyValuePairs feature.

Generative AI

AI that creates new content (text, images, audio, code) from a prompt; predicting a number or a class is predictive ML, not generative AI.

GitHub Copilot

AI pair programmer that suggests code and can explain code and add inline comments or documentation.

Global Standard (GlobalStandard)

Pay-per-token deployment type routed across Azure's global infrastructure with the highest default quota; the recommended starting point.

GPT (Generative Pretrained Transformer)

OpenAI's family of transformer LLMs pretrained on large text corpora for understanding and generating language and code.

GPT-3.5 Turbo

Earlier text-in, text-out OpenAI chat model; it can't interpret images or transcribe speech.

GPT-4.1

OpenAI chat model series (gpt-4.1, mini, nano) with text and image input and text output; now deprecated in Azure (gpt-4.1-nano retires on 14 October 2026, gpt-4.1 and mini on 14 April 2027).

GPT-4o

Multimodal OpenAI chat model that accepts text and images (and audio in some variants); supports temperature and penalty parameters.

gpt-4o-transcribe

Speech-to-text model powered by GPT-4o for file-based transcription through the Audio API; its 2025-03-20 version retires on 15 October 2026.

GPT-5 series

OpenAI reasoning model series (gpt-5, mini, nano and later) with text and image input; doesn't support temperature or penalty parameters.

gpt-image-1 (gpt-image series)

OpenAI image generation model series that creates images from text prompts; it replaced DALL-E in Azure. The original gpt-image-1 (preview) itself retires on 23 October 2026, so new work uses gpt-image-1.5, gpt-image-1-mini or gpt-image-2.

Groundedness evaluator

RAG evaluator that measures whether a response stays consistent with the provided context without fabricating content (the precision aspect).

Grounding (grounding data)

Supplying relevant, current facts in the prompt so the model answers from that context, reducing fabrication and enabling citations; the core of RAG.

Grounding data

Trusted content supplied in the prompt at request time so the model answers from it rather than guessing; it reduces fabrication.

Guardrail intervention point

Point where a Foundry guardrail control scans content: user input, tool call (preview, agents only), tool response (preview, agents only) or output.

H

Harm categories

The four classified content risks, hate and fairness, sexual, violence and self-harm, each scored safe, low, medium or high for text and images.

Hate and unfairness evaluator

Safety evaluator that detects hateful or unfair content in model or agent outputs; one of the risk and safety evaluators alongside violence, sexual, self-harm and protected material.

Hosted agent

Foundry agent you build in your own code and framework, which Foundry runs as a container with a managed endpoint; contrast with a declarative prompt agent.

Human in the loop

Design where a person reviews, approves or overrides AI decisions; the exam maps it to the accountability principle.

I

ID document model (prebuilt-idDocument)

Document Intelligence prebuilt model that extracts fields such as name, date of birth, document number and expiry from driver licences and passports.

Image Analysis

Azure Vision API that extracts visual features from an image in one call, such as a caption, Read OCR text, tags and detected objects with bounding boxes.

Image Analysis 3.2

Older Image Analysis version with the widest feature set: tags, objects, Describe captions, brands, faces, image type, colour scheme, landmarks, celebrities and adult content.

Image Analysis 4.0

Newer Image Analysis version with better models for Read, Caption, dense captions, tags, objects, people and smart crop; deprecated and retiring 25 September 2028.

Image captioning

Azure Vision feature that generates a one-sentence, human-readable description of an image (dense captions describe up to 10 regions).

Image classification

Computer vision technique that assigns a label (or labels) to the whole image with no location, for example bear species or benign vs malignant.

Image detail level (detail)

Input_image setting (low, high or auto default) that trades resolution and token use; low processes a 512x512 version.

Image generation tool (image_generation)

Built-in Foundry Agent Service tool that generates images from text prompts using a gpt-image-1 deployment plus an orchestrator model; it's a tool, not an input content type.

Image tagging (tags)

Image Analysis feature that returns one-word tags for objects, living things, scenery and actions, each with a confidence score.

image_url

Field of an input_image item holding a fully qualified public URL or a base64 data URI; the URL must be reachable from Microsoft infrastructure, so private or firewall-restricted URLs fail.

Images API (image generation API)

Azure OpenAI endpoint (images/generations) that creates images from prompts with gpt-image models; use it directly for editing and masks.

Immersive Reader (Azure AI Immersive Reader)

Foundry Tool that improves reading comprehension by isolating text, reading it aloud, translating and highlighting parts of speech.

Inclusiveness principle

Responsible AI principle that AI empowers and engages everyone, removing barriers through accessibility, language support and screen-reader or assistive technology.

Inference

Using a trained model to generate a response; the deployed model isn't retrained per request and doesn't look up stored training documents.

Information extraction

Workload that pulls structured fields from documents, images, audio and video, such as forms, invoices and receipts.

input_image

Responses API content item for an image; its image_url is a public URL or base64 data URI (or file_id for a Files API upload), with optional detail. The type is input_image in every case.

input_text

Responses API content item carrying the prompt text (text field) alongside input_image items in one user message.

Instant access (instant models)

Foundry preview that lets you call supported models by name in the playground or code with no deployment; deploy instead when you need reserved throughput, custom content filters or data residency.

Intent recognition (Speech) (IntentRecognizer)

Former Speech SDK feature that mapped utterances to intents; retired 30 September 2025, so use speech to text followed by CLU or an Azure OpenAI model.

Invoice model (prebuilt-invoice)

Document Intelligence prebuilt model that extracts invoice fields such as InvoiceId, dates, vendor, customer, line items and totals.

J

JSON (JavaScript Object Notation)

Text data format; Azure OpenAI REST requests and responses use application/json bodies.

JSONL (JSON Lines)

Format with one JSON object per line, required for fine-tuning training data (chat format, UTF-8 with BOM, under 512 MB) and batch files; not used for normal REST calls.

K

Key phrase extraction

Azure Language feature returning a plain list of the main concepts in text with no categories or confidence scores.

Keyframe (key frame)

Representative frame sampled from each video shot that Content Understanding uses as visual input for field extraction.

Keyless authentication (Microsoft Entra ID authentication)

Recommended way to call Foundry models with a Microsoft Entra ID token and an RBAC role (such as Foundry User on the project) instead of an API key.

Knowledge mining

Indexing and AI-enriching large volumes of content so it can be searched and analysed, typically with Azure AI Search.

L

Language detection

Azure Language feature that identifies the language a text is written in; it doesn't translate or filter.

Language identification (LID)

Speech feature that identifies which of up to 4 (at-start) or 10 (continuous) candidate languages is spoken in audio, used with speech to text or speech translation; it doesn't filter profanity.

Language Studio

Legacy web portal for trying and customising Azure Language features; not a model catalog, and Language features are now used in the Foundry portal.

Layout model (prebuilt-layout)

Document Intelligence model that extracts text, tables, selection marks and document structure, plus key-value pairs with the keyValuePairs feature.

Limited Access

Microsoft policy requiring registration and approval (managed customers, approved use cases) before using sensitive features such as Speaker Recognition, Face identify/verify and custom neural voice.

LLM (large language model)

Neural network (usually a transformer) with billions of parameters trained on massive text to predict the next token; broadly capable but costlier and slower than an SLM.

LROPoller (long-running operation poller)

Azure SDK object for a long-running operation: result() polls to completion and returns the result, status() only reports state, wait() blocks without returning it.

LUIS (Language Understanding)

Legacy intent service replaced by CLU. Historically scheduled to retire on 1 October 2025; today its runtime and authoring endpoints are fully retired (since 31 March 2026) and all LUIS requests fail.

M

Machine translation

Automatic conversion of text from one language to another by a model such as Translator or a generative model; language modelling is how models learn, not a user-facing conversion.

Managed compute

Foundry deployment option that places model weights on dedicated, Foundry-managed GPU virtual machines behind a REST endpoint, billed per compute hour; used for open-source and custom models.

Markdown

Lightweight text format Content Understanding uses to represent extracted content (text, tables, transcripts) alongside fields.

Max completion tokens (max_completion_tokens, max response)

Upper bound on generated tokens, including reasoning tokens, so it caps response length and cost.

max_tokens (max completion tokens, max response)

Parameter that caps how many tokens the model generates in a response; input plus output must fit the model's context length.

Metaprompt and grounding mitigation layer

Application-level generative AI mitigation layer where the system message and grounding data steer the model away from harmful or ungrounded output.

Microsoft Entra ID (Azure AD)

Microsoft's cloud identity service and tenant for Azure and Microsoft 365.

Microsoft Foundry (Azure AI Foundry)

Azure platform for building, evaluating and running AI agents and models under one resource, with guardrails, evaluations and AI red teaming.

Microsoft Foundry resource (Foundry resource)

Top-level Azure resource (Microsoft.CognitiveServices) that holds model deployments, security and networking settings and child projects; required for Content Understanding, unlike a standalone Azure OpenAI resource.

Microsoft Responsible AI Standard

Microsoft framework for building AI systems on six principles: fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability.

ML (machine learning)

Training models that learn patterns from data to make predictions such as regression and classification; predictive ML is not generative AI.

Model catalog

Foundry portal hub to discover, compare, test and deploy models from Microsoft, OpenAI, Anthropic, Meta, Mistral, Cohere, Hugging Face and others, filtered by task, capability and deployment option.

Model deployment (deployment)

A named instance of a Foundry model in a resource, with a deployment type and TPM quota; the deployment name is what requests pass as the model.

Model deployment name

Name you give a deployment and pass in the model parameter; requests route by deployment name (for example my-mini-gpt), not the underlying model name.

Model leaderboard

Foundry model catalog view (preview) that ranks models on quality, safety, estimated cost and throughput benchmarks, with trade-off charts and side-by-side compare of up to three models.

Model mitigation layer

Lowest generative AI mitigation layer: choosing and understanding (or fine-tuning) the model; changing model is not a safety control on its own.

Model parameters (weights)

Numeric values learned during training that encode a model's knowledge; their count is what separates an SLM from an LLM.

Model temperature (temperature)

Sampling parameter (0 to 2) that sets output randomness: low values give focused, repeatable answers, high values more creative ones; adjust it or Top P, not both.

Model version upgrade policy

Deployment setting (auto-update to default, upgrade when expired, or no auto upgrade) that controls when a Standard deployment moves to a new model version.

Multi-service Foundry resource (Microsoft Foundry resource)

Single Azure resource giving one endpoint and key for Azure OpenAI and multiple Foundry Tools, as opposed to a single-service resource for one tool.

Multimodal model (vision-enabled model)

Generative model that accepts more than one input type, such as text plus images, in a single request.

N

NER (named entity recognition)

Azure Language feature that finds entities and labels them with categories such as Person, PersonType, Location, Organization, DateTime, Quantity and Skill.

Neural voice (standard voice)

Prebuilt deep-neural-network text to speech voice available out of the box in 100+ languages and locales.

Next-token prediction

Transformer stage that computes a probability distribution over the vocabulary, picks a token, appends it and repeats to build output.

NLP (natural language processing)

Workload that interprets existing written language, such as sentiment, key phrases, entities and classification.

O

o-series

OpenAI reasoning models (o1, o3, o4-mini) that spend more time reasoning for science, code and maths tasks; all are deprecated in Azure and retire on 19 November 2026, with GPT-5.6 models as replacements.

Object detection

Computer vision technique that returns a label, confidence score and bounding-box coordinates for each object instance, so you can locate and count multiple objects or classes.

OCR (optical character recognition)

Extracting printed or handwritten text from images and documents so it can be analysed; Azure Language can't read scanned images itself.

Opinion mining (aspect-based sentiment analysis)

Sentiment analysis option that links sentiment to specific targets (aspects) and their assessments in the text.

output_text

Convenience property of a Responses API result that holds the model's generated text; output items are responses, not input content types.

P

Past messages included (conversation history)

Playground setting for how many previous chat turns are sent with each request; more history gives context but uses more tokens.

PII detection (personally identifiable information detection)

Azure Language prebuilt feature that detects entities such as Person, PhoneNumber, Email and Address and returns redacted (masked) text; not harmful-content moderation.

Prebuilt analyzer

Ready-to-use Content Understanding analyzer such as prebuilt-read, prebuilt-layout, prebuilt-invoice, prebuilt-audioSearch or prebuilt-callCenter; can be customised with extra fields.

prebuilt-audio

Content Understanding base analyzer for audio that transcribes to WebVTT with diarization before field extraction.

prebuilt-audioSearch

Prebuilt Content Understanding audio analyzer that returns a diarized transcript and a one-paragraph conversation summary.

prebuilt-callCenter

Prebuilt Content Understanding post-call analyzer that returns a transcript with speaker roles plus summary, sentiment, topics, companies, people and categories.

prebuilt-document

Content Understanding base analyzer for documents (PDFs including scanned, Office files, forms) that the OCR and layout analyzers build on.

prebuilt-image

Content Understanding base analyzer for still images; it doesn't take PDFs (use a document analyzer).

prebuilt-invoice

Domain-specific Content Understanding analyzer that extracts invoice fields including line items as structured data.

prebuilt-layout

Content Understanding content extraction analyzer that adds layout (tables, figures, sections, paragraphs) to OCR; no language model needed.

prebuilt-read

Content Understanding content extraction analyzer that OCRs words, paragraphs, formulas and barcodes without layout; synchronous analysis is preview only.

prebuilt-video

Content Understanding base analyzer for video that extracts transcript, keyframes and shots and supports segmentation into scenes.

Presence penalty (presence_penalty)

Inference parameter (-2.0 to 2.0) that penalises any token already present; positive values push the model to new topics.

Privacy and security principle

Responsible AI principle that personal and business data is protected, with notice, consent and restricted access; it is not about handling unusual inputs.

Prompt agent

Declaratively defined Foundry agent made of a model, instructions and tools that Foundry runs for you with no code or containers to manage.

Prompt engineering

Crafting prompts, system messages and examples to steer a model's output without changing its weights.

Protected material evaluator

Risk and safety evaluator that detects copyrighted or otherwise protected text (such as lyrics or articles) in outputs.

Provisioned throughput (PTU, provisioned throughput units)

Deployment type that reserves dedicated model capacity measured in PTUs and billed hourly whether used or not, for predictable throughput and latency.

R

RAG (retrieval augmented generation)

Pattern that retrieves relevant content from your data (often via an index), adds it to the prompt as grounding data and generates a cited answer; used for private or recent information.

Read (OCR) (Read API)

Azure Vision OCR engine that extracts printed and handwritten text with locations and confidence scores; for text-heavy documents use Document Intelligence or Content Understanding.

Read model (prebuilt-read)

Document Intelligence OCR model that extracts printed and handwritten text, lines, words and languages only; recommended for text-heavy scans and the engine under the other models.

Real-time transcription

Speech to text mode that transcribes streaming audio (microphone or file) as it is recognised, returning intermediate results; used for live captions and voice input.

Reasoning model

Model (GPT-5 series, o-series) that generates hidden reasoning tokens before answering; it doesn't support temperature, top_p or penalty parameters and uses max_completion_tokens.

Receipt model (prebuilt-receipt)

Document Intelligence prebuilt model that extracts sales receipt fields such as MerchantName, transaction date, tax and total.

recognize_once() (recognize_once_async())

SpeechRecognizer single-shot method that returns one utterance, ending at silence or after a maximum duration (about 15 to 30 seconds); recognize_once_async().get() is the non-blocking form. Use for short commands.

Regression

Supervised ML that predicts a numeric value, such as revenue or rentals; predictive, not generative.

Relevance evaluator

RAG evaluator that measures how accurately and directly a response addresses the user's query.

Reliability and safety principle

Responsible AI principle that AI performs as intended across conditions, responds safely to unusual or missing input and resists manipulation; declining to predict on missing fields belongs here.

Responses API

Azure OpenAI and Foundry API that takes an input array of messages and content items (input_text, input_image) and returns output items, with the text in response.output_text.

Responsible AI (RAI)

Approach to developing, assessing and deploying AI systems safely, ethically and with trust; it continues after deployment through production monitoring.

Responsible AI dashboard

Azure Machine Learning interface combining error analysis, fairness assessment, interpretability, counterfactual and causal analysis to debug models.

RMSE (root mean squared error)

Regression metric summarising prediction error as the square root of the mean squared difference between predicted and actual values; lower is better and it doesn't apply to classification.

S

Safety system message

System message that adds explicit boundaries and refusal guidance to mitigate harms; one layer of a safety stack, not a complete control.

Safety system mitigation layer

Platform-level generative AI mitigation layer of guardrails (content filters) and abuse monitoring that classify and block harmful prompts and completions.

SAS

See Shared access signature.

SDK (software development kit)

Client libraries, such as the Speech SDK or the Content Understanding client library, that wrap a service's REST API.

Self-supervised learning

Training that derives labels from the data itself, such as predicting the next token, used to pretrain foundation models; not supervised learning on labelled data.

Semantic segmentation

Computer vision technique that classifies every pixel to produce a mask of an object's exact shape (Azure Machine Learning AutoML supports the related instance segmentation task); goes beyond a bounding box.

Sentiment analysis

Azure Language feature returning positive, neutral, negative or mixed labels with confidence scores per sentence and document.

Serverless API deployment

Preferred Foundry deployment option where Microsoft hosts the model on shared infrastructure and you pay per token (or PTU), with deployment types such as Global Standard, Standard and Provisioned.

Severity threshold

Guardrail setting per harm category (default medium) at or above which content is blocked; blocking at low on input and output is the strictest setting.

Shared access signature (SAS)

Signed token granting time-limited delegated storage access; a signature, not a role assignment, and not applicable to SMB.

SLM (small language model)

Language model with the same kind of architecture as an LLM but fewer parameters (roughly under 10 billion), so cheaper and faster but less broadly capable.

speak_text_async()

SpeechSynthesizer method that synthesises plain text without blocking and returns a future for the SpeechSynthesisResult; speak_ssml_async() takes SSML instead.

Speaker recognition (voice biometrics)

Azure Speech capability that identifies or verifies who is speaking from voice characteristics; a Limited Access feature, distinct from speech recognition, which transcribes what is said.

Speaker role detection

Content Understanding audio feature that maps diarized speakers to roles such as agent and customer in call-centre recordings.

Speech captioning (captioning)

Speech to text scenario that produces time-coded captions (SRT or WebVTT) for live or prerecorded audio, with profanity filter options.

Speech profanity filter (profanity option)

Speech to text setting that masks (default), removes or shows (raw) profane words in transcripts and captions; it's a Speech setting, not language detection.

Speech recognition (speech to text, STT)

Converting spoken audio into text, such as captions and call transcripts; a separate workload from NLP on AI-901.

Speech SDK

Client library (Python package azure-cognitiveservices-speech, imported as speechsdk) exposing SpeechConfig, SpeechRecognizer, SpeechSynthesizer and audio config classes.

Speech synthesis (text to speech, TTS)

Converting text into natural-sounding spoken audio, such as reading messages aloud.

Speech to text (speech recognition)

Azure Speech capability that converts spoken audio into text (captions, call and meeting transcripts, voice commands); it doesn't identify who is speaking.

Speech translation

Azure Speech capability that translates spoken audio in real time into text or synthesised speech in one or more target languages.

SpeechRecognizer

Speech SDK class that performs speech recognition from a microphone, file or stream and returns transcribed text, singly or continuously.

SpeechSynthesizer

Speech SDK class that converts text to speech, taking a SpeechConfig and an AudioOutputConfig; speak_text_async() belongs here, not to SpeechRecognizer.

SSML (Speech Synthesis Markup Language)

XML-based markup for text to speech that controls voice, pitch, rate, volume, pauses and pronunciation; sent with speak_ssml_async().

Standard deployment type (Standard)

Pay-per-token Foundry deployment type that processes data within the resource's Azure geography, for geography compliance at lower volume.

start_continuous_recognition()

SpeechRecognizer method for long, multi-utterance audio that returns results through recognizing, recognized and canceled event handlers until stop_continuous_recognition() is called.

start_keyword_recognition() (keyword recognition)

SpeechRecognizer method that listens locally for a wake word from a keyword model and then starts sending audio to the service; not general transcription.

Stop sequence (stop)

Up to four strings at which the model stops generating; the output ends before the sequence.

Summarization

Azure Language feature producing extractive (key sentences) or abstractive (newly worded) summaries of text, conversations or documents.

Supervised learning

Machine learning that trains on labelled examples to predict a known output, such as classification or regression.

Synapse workspace (Azure Synapse Analytics)

Analytics service combining SQL and Spark pools and pipelines in a workspace; with a managed VNet it reaches data stores only through managed private endpoints.

System message (system prompt, metaprompt)

High-priority instructions and context sent first in a chat request to set the model's role, tone, format and boundaries; it steers but doesn't guarantee compliance.

T

Task completion evaluator

Agent evaluator (preview) that measures whether an agent completed the requested task end to end.

Temperature

Inference parameter (0 to 2) controlling randomness; low values such as 0.2 give focused, near-deterministic output and high values more creative output.

Text Analytics

See Azure Language in Foundry Tools.

Text to speech (speech synthesis)

Azure Speech capability that converts text into humanlike spoken audio with standard or custom neural voices, tuned with SSML; the opposite direction to speech recognition.

Time-series forecasting (forecasting)

Predicting future values such as sales or demand from historical time data; not a generative or NLP workload.

Token

Chunk of text (word, part word or punctuation) that an LLM processes; context windows, limits and billing are measured in tokens.

Tokenization

First transformer stage that splits text into tokens and maps each to a vocabulary ID.

Tool call accuracy evaluator

Agent evaluator that measures whether an agent chose the right tools and passed correct parameters without redundant calls.

tool_choice

Request parameter that controls tool calling: auto (default, model decides), required (must call at least one tool), none (no tool calls) or a specific tool's specification to force that tool.

Top P (nucleus sampling)

Inference parameter that limits token choice to the top probability mass, controlling diversity rather than length; adjust it or temperature, not both.

top_p (Top P, nucleus sampling)

Sampling parameter that limits token choice to the top probability mass (0.1 means the top 10%); an alternative to temperature.

TPM (tokens per minute)

Unit of Azure OpenAI quota and deployment rate limit; exceeding it returns HTTP 429.

Trade-off chart

Model leaderboard chart that plots models on two metrics, such as quality against estimated cost, to show which balance both.

Training cutoff

Date after which a model has no knowledge; recent events need grounding at request time (RAG) rather than retraining.

Transformer

Neural network architecture behind GPT and most LLMs: tokenization, embedding, attention over context, then next-token prediction.

Translator detect operation (/detect)

Translator operation that returns a text's language code and confidence score without translating it.

Transparency principle

Responsible AI principle that people know when AI is used, what it can and can't do and how it reaches outputs, such as explaining a loan decision.

U

User experience mitigation layer

Generative AI mitigation layer of UI design, AI disclosure, review-and-edit prompts and input/output limits that prevent misuse and overreliance.

User message (user prompt)

Chat message role that carries the end user's question and inputs such as an image, as opposed to the system message's standing instructions.

V

Vision-enabled chat model (large multimodal model)

Chat model (GPT-4o, GPT-4.1, GPT-5 series, o-series) that interprets images in the prompt; a text-only model can't be given this ability by the app layer.

Voice Live API (Voice Live)

Fully managed, low-latency speech-to-speech API for real-time voice agents that combines speech recognition, a generative model and text to speech in one WebSocket interface, with optional avatar and function calling; it returns audio, not only text.

Voice-based prompt agent

Foundry prompt agent configured with a model, instructions, audio settings and tools that you talk to over Voice Live instead of hosting voice orchestration yourself.

W

WebVTT (Web Video Text Tracks)

Timed-text format (hh:mm:ss.fff) used for Speech captions and Content Understanding audio and video transcripts, with speaker tags.

Whisper

OpenAI speech-to-text model for transcription and translation into English (25 MB file limit); available in Azure OpenAI and Azure Speech.