In a transformer, the mechanism determining the weight each other token in the context carries when predicting the next one.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Attention in context, with comparison tables and the common traps.
Terms in this definition
- Transformer
The neural network design underlying GPT and most other LLMs. Text is tokenised and embedded, attention is applied across the context, and the next token is predicted.
- Token
The unit of text an LLM works with, which may be a word, part of a word or punctuation. Billing, limits and context windows are all counted in these units.
- NEXT
Restricted to visual calculations, this DAX function reads the value one step further along an axis of the visual matrix; writing OFFSET with 1 gives the same result.
Related terms
- Alert Triage Agent
A preview Security Copilot agent that sorts DLP alerts for you, weighing the risk each one carries and placing it in groups like Needs attention or Less urgent. Running it uses security compute units.