Which model setting controls randomness, which controls length and cost, and which ones are not set at deployment.
From Ultra Transcenders AI-901 by Tony Rough (publishing soon)
Once a model is deployed, you can shape each response by setting parameters in your code or in the playground. These control how random, how long and how repetitive the output is.
| Parameter | Controls |
|---|---|
| Temperature | Randomness; low (e.g. 0.2) = focused, consistent, near-deterministic; high = varied, creative |
| Top P (nucleus sampling) | Diversity of token selection, not length |
| Max completion tokens / max response | Upper bound on generated tokens (including reasoning tokens), so caps length and cost (output is billed per token) |
| Presence penalty (-2.0 to 2.0) | Penalises any token already seen; positive values push to new tokens and topics |
| Frequency penalty | Scales with how often a token appeared; mainly reduces verbatim repetition |
| Stop sequence | Ends generation at a chosen string |
| Past messages included | How much chat history is sent |
Common trap: Top P controls how diverse the token choices are, not the length of the response. To limit length (and cost), set max completion tokens.
It’s important to know which settings belong to each request and which belong to the deployment.
Common trap: Temperature is not part of the deployment configuration; you set it on each request.
Not every model accepts every parameter. Reasoning models (GPT-5 series, o-series) don’t support temperature, Top P or the penalty parameters; chat models such as GPT-3.5 Turbo and GPT-4o do.
This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-901 glossary · All AI-901 study notes
Fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability, and how to tell them apart in a scenario.
How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.
What happens between a prompt and a response, and how inference differs from training.
How to recognise each AI workload from a scenario.
Agents as model plus instructions, knowledge and tools, and the auto, required and none tool_choice values.
Speech to text, text to speech, translation, batch transcription and speaker recognition compared.
How analyzers turn documents, images, audio and video into structured JSON, and how to call them from code.