A setting used at inference, ranging from -2.0 to 2.0, that lowers a token's likelihood according to how many times it has already occurred. Its main effect is to cut down word-for-word repetition.
Also called frequency_penalty.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Frequency penalty in context, with comparison tables and the common traps.
Terms in this definition
- Inference
The act of producing a response with a trained model. Each request neither retrains the deployed model nor makes it look up stored training documents.
- Token
The unit of text an LLM works with, which may be a word, part of a word or punctuation. Billing, limits and context windows are all counted in these units.