The act of producing a response with a trained model. Each request neither retrains the deployed model nor makes it look up stored training documents.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Inference in context, with comparison tables and the common traps.
Related terms
- az ml online-deployment get-logs
Azure Machine Learning CLI command fetching container logs for a deployment, from the inference server by default or from storage-initializer, to troubleshoot issues like init() failures or missing packages.
- Cognitive Services User
Built-in role permitted to call any model or Foundry Tool on the resource and to list and read its keys, which is wider access than Azure OpenAI inference alone requires.
- Deployment type (Foundry Models)
Chosen when you deploy a model, it decides two things: billing (provisioned, batch or standard pay-per-token) and where inference data may be processed (anywhere globally, within a data zone or inside one geography).
- Frequency penalty
A setting used at inference, ranging from -2.0 to 2.0, that lowers a token's likelihood according to how many times it has already occurred. Its main effect is to cut down word-for-word repetition.
- Global Provisioned
Deployment type that reserves PTU capacity and allows inference to run in any Azure region. Processing is therefore not kept within the resource's geography, as it would be with Standard.
- Model monitoring
Capability in Azure Machine Learning that, on a schedule, checks production inference data against reference data for each signal. When a metric crosses its threshold it raises an alert, by email unless Event Grid events are configured.
- Presence penalty
Setting between -2.0 and 2.0 applied at inference that penalises tokens which have already appeared at all. Raising it above zero encourages the model to move on to new topics.
- RequestResponse log
Log category on Azure OpenAI resources that records each request with its latency and status code. Administrative operations go to Audit logs, and detailed inference traces to Trace logs.