Sampling setting between 0 and 2 controlling how random the output is. Lower values yield focused, repeatable responses and higher ones more creative text; tune this or Top P, but not both.
Also called temperature.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Model temperature in context, with comparison tables and the common traps.
Terms in this definition
- VALUES
Returns in DAX the distinct column values, or table rows, still visible after filters are applied, sometimes with an extra blank entry. CALCULATE often takes the result as a table filter.
- Top P
Also known as nucleus sampling, this inference parameter restricts token selection to the most probable share of the distribution; 0.1 keeps the top 10%. It shapes diversity, not length, and should be tuned instead of temperature, not with it.
Related terms
- Chat playground
Lets you try a deployed chat model in Foundry without writing code, adjusting the system message, temperature, past messages included and maximum response length.
- Claude models in Foundry
Anthropic models available to deploy through Microsoft Foundry. From Opus 4.7 onwards they support adaptive thinking only, refusing temperature and an enabled thinking type, whereas Sonnet 4.6 still takes both.
- Foundry playground
Interactive space in the Foundry portal where a deployed model can be tried out before any code is written. Users tweak temperature, top_p, max tokens and the system prompt, compare models and export sample code; production traffic does not belong here.
- GPT-4o
OpenAI chat model that is multimodal, taking images and text and, in some variants, audio. Temperature and penalty parameters work with it, and the 2024-08-06 version allows supervised, DPO and vision fine-tuning.
- GPT-5 series
A series of OpenAI reasoning models, including gpt-5, mini and nano and their successors, that take text and images. Penalty and temperature parameters are not supported.
- Reasoning model
Kind of model, for instance the o-series or GPT-5 series, that produces hidden reasoning tokens before its answer. Temperature, top_p and penalty parameters are unsupported, and
max_completion_tokenssets output length. - top_p
Also known as nucleus sampling, this inference parameter restricts token selection to the most probable share of the distribution; 0.1 keeps the top 10%. It shapes diversity, not length, and should be tuned instead of temperature, not with it.