#
${{inputs.<name>}} (job input expression)
Expression in a command that Azure Machine Learning replaces at run time with the input's mounted or downloaded path (or literal value); the flag in front must match the script's argparse name.
A
abfss:// (Azure Blob File System (secure))
URI scheme for ADLS Gen2 over the dfs endpoint (account.dfs.core.windows.net); it doesn't address the blob endpoint.
Action groups
Azure Monitor notification targets; they deliver alerts but do not detect problems themselves.
ADLS Gen2 (Azure Data Lake Storage Gen2)
Standard GPv2 storage with hierarchical namespace enabled, giving real directories and POSIX-style ACLs for analytics.
Agent endpoint (stable agent endpoint)
Stable URL that consumers call by agent name; its version selector either uses "Always use latest" (new versions served with no redeploy) or pins one active version.
Agent Monitoring Dashboard
Foundry monitoring view backed by Application Insights showing token usage, latency, run success rate and evaluation scores, with alerts; runtime visibility, not a release gate.
Agent version
Immutable snapshot of an agent's configuration; any change, even one prompt edit, creates a new version, and rollback means pointing at a prior version.
AKS (Azure Kubernetes Service)
Managed Kubernetes with full cluster and node-pool control; autoscales with HPA and the cluster autoscaler; has no built-in user sign-in.
Alert rule
Azure Monitor rule combining a scope (target resources), a condition (signal and logic) and optional action groups; each distinct signal with its own recipients needs its own rule.
AmlOnlineEndpointConsoleLog
Log Analytics table of the console output written by online endpoint containers, used to debug containers that fail to start or mishandle requests.
API key (Foundry) (key-based authentication)
Shared secret for calling a Foundry or Azure OpenAI endpoint that grants full access with no role checks; sharing it removes role separation, and it gives no network isolation.
App registration
Entra ID object defining an application's identity, permissions and supported account types; used for OpenID Connect sign-in and multi-tenant apps.
Application Insights
Azure Monitor APM service for app telemetry, Application Map, availability tests and usage analytics; workspace-based instances store data in Log Analytics.
argparse
Python standard-library module a training script uses to read named command-line arguments; a flag that doesn't match the declared name errors or leaves the argument unset.
Autoscale
Azure Monitor feature that scales out or in (adds or removes instances) when a metric rule holds for its whole duration or on a schedule; the maximum instance count is only a cap and never forces scaling.
az ml model restore
CLI command that brings an archived model or version back into list queries, for audit or rollback.
az ml online-deployment get-logs
CLI command that returns a deployment's container logs (inference server by default, or storage-initializer) to diagnose missing packages or init() failures.
az ml online-endpoint update --traffic
CLI command that sets the live traffic split between an endpoint's deployments (for example "blue=90 green=10"); the percentages must total 100 or 0.
Azure AI Administrator
Built-in role with all control plane permissions for Azure AI and its dependencies; applies to Azure Machine Learning and Foundry hubs only, so current Foundry resources use Foundry Account Owner or Foundry Owner.
Azure AI Developer
Built-in role for all actions inside an Azure Machine Learning workspace or hub except managing it; not for Foundry projects, which use Foundry User or Foundry Owner.
Azure AI Search
Azure search service that indexes content for full-text, vector, hybrid and semantic-ranked retrieval, used as the retrieval layer for RAG and Foundry agents.
Azure CLI (CLI)
Cross-platform command line, e.g. az policy state trigger-scan or az storage queue.
Azure Container Instances (ACI)
Per-second-billed container groups with no cluster; in Azure Machine Learning a legacy (v1) deployment target with no autoscale and no low-priority VMs.
Azure Firewall
Managed stateful network firewall, deployable in Virtual WAN hubs and managed by Firewall Manager.
Azure Login action (azure/login)
GitHub Actions step that signs the workflow in to Azure as an Entra app's service principal or a user-assigned managed identity, ideally via OpenID Connect so no secret is stored.
Azure Machine Learning (AML)
Azure cloud service for training, deploying, monitoring and retraining ML models (MLOps), including online endpoints, pipelines, model monitoring and prompt flow.
Azure Machine Learning CLI v2 (az ml)
Azure CLI ml extension (az ml ...) for creating jobs, assets, endpoints and deployments from YAML.
Azure Machine Learning component (component)
Self-contained, versioned, reusable pipeline step with a name, inputs, outputs, command, code and environment, analogous to a function.
Azure Machine Learning environment (environment)
Versioned asset that packages a Docker image plus a conda or pip specification so jobs and deployments run with reproducible dependencies; reference it by name and version or @latest.
Azure Machine Learning pipeline (ML pipeline)
Workflow that chains components (for example prepare data, train, score) into a reusable, schedulable job.
Azure Machine Learning registry (Azure ML registry)
Central store of versioned models, environments, components and data assets shared by many workspaces (dev, test, prod) with lineage and its own access control; it complements workspaces rather than replacing them.
Azure Machine Learning SDK v2 (azure-ai-ml)
Current Python SDK (MLClient) for Azure Machine Learning; it has no logging API of its own, so metrics are logged with MLflow.
Azure Machine Learning workspace (AML workspace)
Top-level Azure Machine Learning resource that holds models, environments, components and data assets, keeps job history and references datastores and compute; assets registered in it belong to it unless placed in a registry.
Azure Monitor
Azure's unified observability service that collects, analyses and alerts on metrics, logs and traces from Azure and hybrid resources.
Azure Monitor alerts
Rules that notify or act when metric, log or activity-log conditions are met; they don't enforce SAS expiry or storage access.
Azure Monitor Metrics
The metrics half of the Azure Monitor data platform, a time-series store of numeric values; platform metrics are collected automatically and kept 93 days.
Azure Monitor workspace
Store for Prometheus metrics only; logs go to a Log Analytics workspace.
Azure OpenAI (Azure OpenAI in Foundry Models)
OpenAI models (GPT, o-series, embeddings, Whisper) offered as Foundry Models sold by Azure through Serverless API deployments.
Azure Private Link (Private Link)
Platform that exposes PaaS services such as Azure Storage on a private endpoint in your VNet over the Microsoft backbone; needs VNet and DNS setup and does not admit a public IP.
Azure RBAC
Azure role-based access control: assigns built-in or custom roles to users, groups and managed identities at a scope; keyless (Microsoft Entra ID) calls to Foundry need a data-plane role on the resource (for example Cognitive Services OpenAI User).
azureml:// URI (datastore URI)
Azure Machine Learning data reference format azureml://datastores/<name>/paths/<path> that points to data through a registered datastore.
B
Batch endpoint
Endpoint that runs asynchronous jobs over data in storage, processing mini-batches in parallel on a compute cluster and writing results to storage; suited to scheduled scoring of millions of records, not low latency.
Blob storage
Object storage for unstructured data such as video and images; block blobs up to about 190.7 TiB.
Blocklist
Custom list of terms or regex patterns (up to 10,000 per list) attached to a content filter to block specific words; changing it is a content-filter change, not monitoring or evaluation.
Blue-green deployment (safe rollout)
Rollout pattern that adds a new (green) deployment to an online endpoint at 0% traffic, tests it, then shifts traffic gradually; rollback is setting traffic back to blue=100.
BM25 (Okapi BM25)
Relevance ranking algorithm for full-text search that scores documents by term frequency, term rarity and document length.
BOM (byte-order mark)
Marker at the start of a text file; fine-tuning files must be UTF-8 with a BOM and each under 512 MB, while ISO-8859-1, ASCII or UTF-16 can fail validation.
Branch protection
Repository rule (GitHub branch protection or Azure Repos branch policies) that blocks direct pushes to a branch such as main and requires pull request reviews before merging.
Built-in evaluators
Foundry catalogue of quality, RAG, similarity, safety and agent evaluators (plus OpenAI graders and custom evaluators) used to score model and agent outputs.
C
Chunk overlap
Text repeated at the boundary of consecutive chunks to keep context; too much (for example one third) duplicates text so top-k returns near-duplicates.
Chunking
Splitting documents into smaller passages before embedding and indexing; Azure AI Search suggests about 512 tokens with roughly 10-25% overlap, and oversized chunks dilute meaning across unrelated sections.
CI/CD (continuous integration and continuous delivery)
Automated build, test and deployment pipelines (GitHub Actions, Azure Pipelines); running one to apply a change is a redeployment.
Coherence evaluator (CoherenceEvaluator)
General-purpose AI-assisted evaluator scoring logical flow and organisation of a response from the query and response, with no external knowledge.
Cohort
User-defined subgroup of data points in the Responsible AI dashboard, used to compare errors and metrics that a single global cohort would hide.
Command job
Azure Machine Learning job that runs one command (such as python script.py) with declared inputs, outputs, environment and compute.
Compute cluster (AmlCompute)
Managed multi-node CPU or GPU cluster that autoscales between min_instances and max_instances per job; set min_instances to 0 with an idle scale-down time to pay only while jobs run.
Conda specification (conda file)
YAML file listing the Python packages that Azure Machine Learning builds on top of a base Docker image to create an environment; installing packages inside the training script is not a versioned alternative.
Connections Active (ConnectionsActive)
Azure Monitor online endpoint metric counting concurrent TCP connections from clients; feature count and dataset size are not runtime metrics.
Content filter (guardrail controls)
Foundry guardrail configuration that classifies prompts and completions by risk (hate, sexual, violence, self-harm, prompt attacks, protected material) and blocks at a severity threshold; not an observability tool.
ContentSafetyEvaluator
Composite safety evaluator covering hate and unfairness, violence, sexual and self-harm severity from the query and response.
Context window
Total tokens (input plus output) a model can process in one request; for example GPT-4.1-mini supports about 1M tokens, Phi-3-mini (now retired) 128k.
CorrelationRemover
Fairlearn preprocessing transformer that removes correlation between sensitive and other training features, so the model must be retrained.
Cosmos DB (Azure Cosmos DB)
Globally distributed NoSQL database with multi-region writes, automatic indexing and under 10 ms latency.
Curated environment
Prebuilt Azure Machine Learning environment hosted in the Microsoft-managed azureml registry, used as is; you don't edit it, you create a custom environment instead.
Custom evaluators
Evaluators you define with Python code or an LLM-as-judge prompt when built-in evaluators don't fit.
D
Data asset
Named, versioned reference to data (uri_file, uri_folder or mltable) that works like a bookmark to a storage path; it doesn't copy the data in a workspace.
Data drift
Model monitoring signal that compares each input feature's distribution in production with a baseline (usually the training data) using metrics such as Jensen-Shannon distance or PSI; a breach can trigger retraining.
Data zone
Microsoft-defined boundary (US, EU or Asia Pacific) spanning multiple regions and countries within which Data Zone deployment types process data.
Data Zone Batch (DataZoneBatch)
Global Batch equivalent that processes batch jobs only within the data zone, with the same 24-hour target turnaround and discount.
Data Zone Standard (DataZoneStandard)
Pay-per-token deployment type that routes traffic anywhere within a Microsoft-defined data zone (US, EU or APAC) with more quota than a geography deployment; a data zone spans many regions, not one.
Datastore
Workspace reference that holds the connection information for an Azure storage service (Blob, Azure Files, ADLS Gen2) so scripts don't embed keys.
davinci-002
Legacy text-only completion base model; not suitable for vision fine-tuning.
Demographic parity
Fairness constraint requiring similar selection rates across sensitive groups, used to mitigate allocation harms.
Deployment type (Foundry Models)
Serverless API setting that fixes where inference data is processed (global, data zone or single geography) and how you pay (standard pay-per-token, provisioned or batch).
Developer deployment type (DeveloperTier)
Deployment type only for evaluating fine-tuned models: no SLA, no data residency guarantee and a 24-hour lifetime; not for production.
Diagnostic setting
Routes a resource's logs and metrics to storage, Log Analytics, Event Hubs or a partner; up to five per resource.
DPO (direct preference optimization)
Fine-tuning method that learns from paired preferred and non-preferred outputs without a reward model; each record has input, preferred_output and non_preferred_output (each with at least one assistant or tool message), and the two outputs must differ.
E
Embeddings (vector embeddings)
Numeric vectors that represent the meaning of text so that similar content sits close together (compared by cosine similarity); changing the vector length or model means re-embedding everything.
Equalized odds
Fairness constraint requiring similar true-positive and false-positive rates across sensitive groups in binary classification.
Error analysis
Responsible AI component that uses a decision tree and heat map to find cohorts with higher error rates than the overall benchmark.
Evaluation dataset
Fixed set of test inputs (query, optionally context and ground truth) run through each model or prompt variant so results are comparable.
Event Grid
Event router that pushes events to handlers; not a store and not a diagnostic-setting destination for Entra.
ExponentiatedGradient
Fairlearn reduction algorithm that wraps an estimator and retrains it on reweighted data to meet demographic parity or equalized odds for binary classification.
F
Fairlearn
Open-source package behind Azure Machine Learning fairness assessment, providing disparity metrics and mitigation algorithms (reduction and post-processing).
Fairness assessment
Responsible AI component (from Fairlearn) that compares performance and selection-rate disparities across groups defined by sensitive features.
Federated identity credential (workload identity federation)
Trust that lets an external workload (e.g. GitHub Actions, Kubernetes) exchange its token for an Entra token without a secret.
Fine-tuning
Further training of a base model on your examples (supervised fine-tuning, DPO or RFT): prepare data, select base model, upload data, create the job, then deploy; the base model needn't be deployed first.
Fluency evaluator (FluencyEvaluator)
General-purpose AI-assisted evaluator scoring grammar, vocabulary, sentence complexity and readability from the response alone.
Foundry Account Owner (Azure AI Account Owner)
Built-in role for managing Foundry resources and projects (deployments, connections, networking) that can assign Foundry User, but has no data actions to build in projects.
Foundry Agent Service
Microsoft Foundry service for building, hosting and running agents (prompt and hosted agents) with tools, versions, tracing and network isolation.
Foundry control plane and data plane
Foundry RBAC split: control plane actions (resource settings, networking, deployments, creating projects) versus data plane actions (building agents, running evaluations, uploading files) inside a project.
Foundry evaluations
Microsoft Foundry scoring of model and agent quality and safety with built-in or custom evaluators; measures risk without blocking at runtime.
Foundry managed virtual network (managed VNet)
Microsoft-managed virtual network that isolates Foundry Agent Service egress and reaches Storage, AI Search and Cosmos DB through managed private endpoints; it can't be disabled once enabled.
Foundry Models sold by Azure
Models (including Azure OpenAI) hosted and billed by Microsoft with Microsoft support, SLAs, Entra ID auth and private networking; the enterprise-governance choice for regulated data.
Foundry Owner (Azure AI Owner)
Highly privileged built-in role with both control plane and data plane permissions: create resources and projects, manage models and build in projects.
Foundry playground (Chat playground)
Foundry portal UI for testing prompts, system messages and parameters against a deployed model before writing code.
Foundry project
Child resource of a Foundry resource that isolates a team's agents, files, evaluations and data while reusing the parent's deployments and connections; governance isn't set here.
Foundry Project Manager (Azure AI Project Manager)
Built-in role that manages and builds in Foundry projects and can assign only the Foundry User role to others.
Foundry User (Azure AI User)
Least-privilege built-in developer role granting reader access plus data actions to build and test in a Foundry project; also assigned to a project's managed identity.
Full-text search (keyword search)
Azure AI Search keyword retrieval over an inverted index ranked with BM25; best for exact identifiers, product codes and names, but misses paraphrases.
G
Git
Distributed version control system; prompts stored in Git get commit author, timestamp, diff, tags and revert.
Git integration (Azure Machine Learning)
Records the repository, branch and commit of the code submitted with each job; it versions code, not registered assets.
GitHub Actions
GitHub workflow automation triggered by events such as push or pull request, used to deploy Bicep or CLI templates and run Azure Machine Learning jobs without a person starting them.
Global Batch (GlobalBatch)
Asynchronous deployment type for large request files with separate quota, 24-hour target turnaround and about 50% lower cost; no real-time SLA, so not for interactive use.
Global Standard (GlobalStandard)
Pay-per-token deployment type that may process data in any Azure region with capacity; highest default quota, newest models first and the recommended default.
gpt-35-turbo (GPT-3.5 Turbo)
Legacy text-only Azure OpenAI chat model; not suitable for vision fine-tuning.
GPT-4.1-mini
Fast, lower-cost Azure OpenAI chat model with text and image input and a context window of about 1M tokens (less on some deployment types); supports SFT and DPO fine-tuning.
gpt-4o (GPT-4o)
Multimodal Azure OpenAI model accepting text and images; version 2024-08-06 supports supervised, DPO and vision fine-tuning.
GridSearch (Fairlearn)
Fairlearn reduction algorithm that retrains an estimator over a grid of reweightings and returns a set of models to choose an accuracy–disparity trade-off from.
Groundedness evaluator (GroundednessEvaluator)
RAG evaluator scoring (1-5) how well a response is supported by the retrieved context; needs context.
H
Hallucination
Model output that states information not supported by its sources or facts; measured indirectly with groundedness, where a lower hallucination rate is an improvement.
Hate and unfairness evaluator (HateUnfairnessEvaluator)
Risk and safety evaluator scoring the severity of hateful or unfair content about individuals or social groups.
Hybrid search
Single Azure AI Search query that runs full-text and vector search in parallel and merges results with Reciprocal Rank Fusion; with semantic ranker it gives the best overall relevance.
I
IaC (infrastructure as code)
Declarative deployment files such as ARM templates and Bicep.
Indirect attack evaluator (XPIA (cross-domain prompt injected attack))
Risk and safety evaluator measuring whether a response fell for a jailbreak injected through retrieved documents or other context.
J
JSONL (JSON Lines)
File format with one JSON object per line, required for fine-tuning data; CSV, TSV or a single JSON array can fail validation.
K
KQL (Kusto Query Language)
Query language of Log Analytics and Azure Data Explorer, used for log alerts.
L
Log Analytics workspace
Store for logs queried with KQL; used by Sentinel, VM insights and workspace-based Application Insights.
Low-priority VM (tier: low_priority)
Historically, cheaper preemptible nodes available only on Azure Machine Learning compute clusters (so batch endpoints could use them); today they are retired (since 31 March 2026), and clusters that specify low_priority still run but get Spot VMs at the Spot rate.
M
Managed compute (Foundry) (managed real-time endpoint)
Foundry deployment option (preview) that hosts open-source, partner and custom models on dedicated GPUs billed hourly; Azure OpenAI models aren't deployed this way.
Managed identity
Entra identity for Azure resources with no stored secret; system-assigned or user-assigned.
Managed online endpoint
Online endpoint whose compute Azure Machine Learning provisions and manages; it supports traffic splitting, mirroring and autoscale, and shares the workspace's private endpoint.
Metric alert
Alert rule that evaluates numeric resource metrics at regular intervals, targeted at the resource; it can't see event log entries.
Microsoft Entra ID (Azure AD)
Microsoft's cloud identity service and tenant for Azure and Microsoft 365.
Microsoft Foundry (Azure AI Foundry)
Azure platform for building, evaluating and running AI agents and models under one resource, with guardrails, evaluations and AI red teaming.
Microsoft Foundry AI agent evaluation GitHub Action (ai-agent-evals)
GitHub Action (preview) that runs offline evaluations in CI/CD: it sends a test dataset to Foundry agents, scores responses with catalogue evaluators, compares versions statistically and can gate a release.
Microsoft Foundry project (Foundry project)
Development boundary inside a Foundry resource for building and evaluating generative AI agents and apps; not a tool for governing traditional ML assets.
Microsoft Foundry resource (Foundry account)
Top-level Azure resource (Microsoft.CognitiveServices/accounts, kind AIServices) where networking, security, policy, model deployments and shared connections are governed; one shared resource with a project per team gives central governance plus isolation.
Mini-batch
Chunk of input files a batch deployment passes to each run() call; mini-batches interrupted by node preemption are rescheduled while completed ones are kept.
MLflow
Open-source tracking and model-management framework that Azure Machine Learning uses as its tracking API; the workspace is an MLflow-compatible tracking server.
MLflow experiment
Named group of runs, set with mlflow.set_experiment, whose parameters and metrics can be compared; it doesn't define a model's artifact path.
MLflow model (mlflow_model)
Model saved in MLflow format (MLmodel file plus conda requirements) that deploys to online and batch endpoints without a scoring script or environment.
MLflow run
One tracked execution (opened by mlflow.start_run or automatically in a job) that records parameters with log_param, metrics with log_metric, and artifacts.
MLflow tracking URI
Address MLflow logs to; set automatically inside Azure Machine Learning jobs and set with mlflow.set_tracking_uri when training elsewhere.
mlflow.autolog
MLflow call that automatically logs parameters, metrics and models for supported frameworks; log_models=False keeps metrics but skips the model.
mlflow.register_model
Registers a logged model from runs:/<run-id>/<artifact-path> into the workspace model registry, keeping lineage to the run.
mlflow.sklearn.log_model (log_model)
Logs a model in MLflow format under the artifact path you pass (for example "model"); the path is exactly that string, not the experiment name.
MLTable (AssetTypes.URI_MLTABLE)
YAML blueprint file (named MLTable) that describes how to read the data files in its folder into a table; used by AutoML and parallel jobs, and loaded by passing the folder, not the file.
mltable Python package (mltable)
SDK for MLTable blueprints: mltable.load() takes the folder containing the MLTable file, to_pandas_dataframe() materialises it, take() samples rows and save() writes a blueprint.
Model archive (az ml model archive)
Hides a model or version from default lists while keeping its files and metadata so it can still be referenced and later restored; unlike delete, it is reversible.
Model interpretability
Responsible AI component (from InterpretML) giving global and local feature importance, showing whether a model relies on sensitive or proxy features.
Model monitoring (Azure Machine Learning model monitoring)
Azure Machine Learning feature that runs a scheduled monitoring job comparing production inference data with reference data per signal, and alerts (email by default, or Event Grid events) when a metric breaches its threshold.
Monitoring signal
One check in an Azure Machine Learning model monitor (data drift, prediction drift, data quality, feature attribution drift, model performance or custom), each with its own metrics and user-defined thresholds.
N
No-code deployment
Deployment of an MLflow model to an online or batch endpoint without a scoring script or environment, which Azure Machine Learning generates.
O
Online deployment
Set of resources (model, environment, scoring script, instance type and count) behind an online endpoint; an endpoint can hold several, such as blue and green.
Online endpoint (real-time endpoint)
HTTPS endpoint for low-latency, synchronous, per-request scoring that can host several deployments behind one scoring URI; managed or Kubernetes.
Online endpoint autoscale
Azure Monitor autoscale rules on an online deployment that add or remove instances by metric (such as CPU or request latency) or schedule; the fix for timeouts that appear only under load.
OpenAI graders (Azure OpenAI graders)
Evaluator type (label model, score model, string check, text similarity) usable in Foundry evaluations and the evaluation GitHub Action.
OpenTelemetry (OTel)
Open standard for traces, metrics and logs that Foundry tracing and Azure Monitor Application Insights use, including GenAI semantic conventions.
P
PF_DISABLE_TRACING
Environment variable controlling prompt flow tracing, which is off by default; set it to false to show the Trace tab with per-node duration and token cost.
Phi-3-mini (Phi-3-mini-128k-instruct)
Small Microsoft language model with at most a 128k-token context (4k variant also existed); now retired in favour of Phi-4-mini-instruct.
Private endpoint
Private IP for a service in your VNet, reachable from peered VNets and from on-premises over ExpressRoute or VPN; public access can then be disabled.
Prompt flow
Azure Machine Learning and Foundry tool for building LLM apps as a graph of nodes (LLM, Python, prompt tools) with variants and evaluation; retires 20 April 2027 in favour of Microsoft Agent Framework.
Prompty (.prompty)
Open Microsoft file format for an LLM prompt (front matter with model and parameters plus a template) that can be stored and versioned in Git.
Protected material evaluator (ProtectedMaterialEvaluator)
Risk and safety evaluator that detects copyrighted text such as song lyrics, recipes and articles in a response.
Public network access
Resource setting (for example on a storage account or an Azure Machine Learning workspace) that governs only the public endpoint; private endpoints still work when it is Disabled.
Pull request (PR)
Request (on GitHub or in Azure Repos) to merge a branch, where reviewers and branch protection rules or policies are applied.
Q
QAEvaluator
Composite evaluator combining groundedness, relevance, coherence, fluency, similarity and F1; needs context and ground truth.
R
RAG (retrieval-augmented generation)
Pattern that retrieves relevant chunks from your content (for example from Azure AI Search) and passes them to an LLM to ground its answer.
RBAC (role-based access control)
Azure role assignments governing who can manage resources, inherited down scopes; not where or what size resources are.
Reference data (model monitoring) (baseline)
The comparison baseline for a monitoring signal, either training/validation data or recent past production data; training data is recommended for data drift.
Registered model (model asset)
Versioned Azure Machine Learning asset (custom_model, mlflow_model or triton_model) with name, version, tags and lineage to the job that produced it; registering again auto-increments the version.
Relevance evaluator (RelevanceEvaluator)
RAG evaluator scoring how well a response answers the query, using only the query and response.
Request Latency (RequestLatency (P50, P90, P95, P99))
Azure Monitor online endpoint metric for the time to respond to a request in milliseconds, also reported at P50 to P99 percentiles.
Requests Per Minute (RequestsPerMinute)
Azure Monitor metric counting requests sent to an online endpoint or deployment each minute; rising with latency indicates saturation.
Responsible AI dashboard (RAI dashboard)
Azure Machine Learning view combining error analysis, fairness assessment, interpretability, data analysis, counterfactuals and causal analysis to check a model before deployment; it doesn't check latency or schema.
runs:/ URI
MLflow URI runs:/<run-id>/<path> pointing to a model in a run's default artifact location, used to register it with lineage; equivalent to azureml://jobs/<job>/outputs/artifacts/paths/<path>.
S
Scoring script (score.py)
Deployment script whose init() loads the model when the container starts and whose run() scores each request; not needed for MLflow models.
Scoring URI
Single URL clients call to invoke an endpoint, which stays the same while deployments behind it change.
Semantic ranker (semantic ranking)
Azure AI Search feature that reranks the top BM25 or hybrid results with Bing-derived language models and can add captions and answers; the re-ranker to add without changing embeddings.
Sensitive feature (sensitive attribute)
Attribute such as age, gender or race that defines groups for fairness assessment and mitigation.
Serverless API deployment (standard deployment)
Foundry deployment option where Microsoft hosts the model and you call it as an API, billed per token or PTU; the path for Foundry Models sold by Azure, including Azure OpenAI, and the one the deployment types apply to.
Serverless compute
On-demand Azure Machine Learning compute used when a job specifies no compute target, so you don't create or manage a cluster and jobs don't queue behind each other.
Service endpoint
Routes a subnet's traffic to a service's public endpoint over the Azure backbone; free, but not usable from on-premises.
Service principal
App-registration identity with a stored secret or certificate that must be rotated and can be copied; suits code outside Azure.
Similarity evaluator (SimilarityEvaluator)
Textual similarity evaluator scoring how closely a response matches an expected answer; needs ground truth.
Similarity threshold (vector search) (minimum similarity score)
Minimum score a vector match must reach to be returned; raising it filters harder and can drop relevant chunks, lowering it admits loosely related ones.
SLA (service level agreement)
Guaranteed uptime, e.g. 99.9% single VM, 99.95% availability set, 99.99% across zones.
Span
Single operation within a trace (for example one prompt flow node or LLM call) with start and end times and attributes; spans nest to show the call sequence.
Spot VMs
Discounted evictable VMs for interruptible work; since 31 March 2026 Azure Machine Learning clusters that request low-priority nodes get Spot VMs instead.
Standard deployment type (regional)
Pay-per-token deployment type that processes prompts and responses within the resource's Azure geography; the option when data must stay in one Azure geography.
start_trace
Prompt flow SDK function that starts tracing to a local trace server and UI unless export is configured, so it doesn't capture traces in the Foundry project by itself.
Storage account
Top-level Azure Storage resource providing a unique namespace for blob, file, queue and table data; its kind, performance and location are fixed at creation.
Synthetic evaluation dataset (data generation service)
Foundry data generation service (preview) output: a versioned dataset of single-turn Q&A pairs or multi-turn simulation seeds generated from an agent definition, prompt or reference file; for pre-launch and regression baselines.
System message
Instruction message at the start of a chat that sets the model's role, rules and format; the system turn in chat and fine-tuning data.
System-assigned managed identity
Identity created and deleted with one resource; ten VMs get ten identities; Azure Policy remediation can use one.
T
text-embedding-ada-002 (ada-002)
Azure OpenAI embedding model that only returns 1,536-dimension vectors (8,192 max input tokens); it generates no text.
ThresholdOptimizer (fairlearn.postprocessing)
Fairlearn post-processing algorithm that applies group-specific thresholds to an existing binary classifier for demographic parity or equalized odds, without retraining.
Time to first token (TTFT)
Latency from prompt submission until the first generated token returns (Azure Monitor Time to Response); larger prompts raise it.
Tokens per second (token generation rate)
Generation throughput of a model deployment; its inverse is Time Between Tokens, which rises when the deployment is under load.
Top-k
Number of nearest-neighbour matches a vector query returns (k); with semantic ranker, Microsoft suggests k of 50 to maximise its inputs.
Tracing (Foundry)
Foundry observability feature (off by default) that records each request as an OpenTelemetry trace with inputs, outputs, tool calls, token usage and latency, stored in the connected Application Insights resource.
Traffic mirroring (shadow testing)
Copies a percentage (up to 50%) of live endpoint traffic to a deployment for testing while clients still get responses only from the live deployment.
U
uri_file (AssetTypes.URI_FILE)
Data asset or job input type that points to a single file of any format.
uri_folder (AssetTypes.URI_FOLDER)
Data asset or job input type that points to a folder or container of many files, mounted or downloaded to the compute.
User-assigned managed identity
Standalone identity attached to many resources, so roles are granted once.
V
Vector search
Similarity search over embeddings that finds content by meaning and paraphrase (nearest neighbours such as HNSW), but can miss exact identifiers.
Virtual network peering (VNet peering)
Private, low-latency connection between VNets in the same or different regions over the Microsoft backbone; on its own it gives App Service no VNet access.
Vision fine-tuning
Fine-tuning with images (URLs or base64) in the JSONL messages; supported only for gpt-4o (2024-08-06) and gpt-4.1 (2025-04-14), with images excluded if they contain people, faces or CAPTCHAs.
VNet injection (Foundry Agent Service) (BYO virtual network)
Network option that injects agent compute into a customer subnet delegated to Microsoft.App/environments (/27 or larger) so traffic to private resources stays in your VNet; set at account creation.
W
wasbs:// (Windows Azure Storage Blob (secure))
URI scheme for Blob Storage over the blob endpoint (account.blob.core.windows.net); use it with uri_folder for all blobs in a container.
weight (fine-tuning)
Optional key (0 or 1) on assistant messages in multi-turn training data; 0 skips training on that message.
Whisper
Azure OpenAI speech-to-text model that transcribes and translates audio; not a text summarisation model.