Home › Glossary › max_tokens

max_tokens

Limits the number of tokens a model produces for a response. The prompt and the output together still have to fit within the model's context length.

Also called max completion tokens, max response.

Read more: Microsoft Learn

In the Ultra Transcenders books

AI-901AI-103

Each book explains max_tokens in context, with comparison tables and the common traps.

See max_tokens in the full glossary