The neural network design underlying GPT and most other LLMs. Text is tokenised and embedded, attention is applied across the context, and the next token is predicted.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Transformer in context, with comparison tables and the common traps.
Terms in this definition
- GPT
A family of large language models from OpenAI, based on the transformer architecture and pretrained on big text collections, used to understand and produce language and code.
- Attention
In a transformer, the mechanism determining the weight each other token in the context carries when predicting the next one.
- NEXT
Restricted to visual calculations, this DAX function reads the value one step further along an axis of the visual matrix; writing OFFSET with 1 gives the same result.
- Token
The unit of text an LLM works with, which may be a word, part of a word or punctuation. Billing, limits and context windows are all counted in these units.
Related terms
- CorrelationRemover
Applied before training, this Fairlearn transformer cancels out the correlation that other features share with sensitive ones; you then need to train the model again.
- LLM
Large language model, usually a transformer network with billions of parameters that learnt to predict the next token from vast amounts of text. Compared with a small language model it is more capable across tasks, but slower and costlier.
- Next-token prediction
Step in a transformer where a probability distribution across the vocabulary is calculated, one token is chosen and added, and the process loops to generate the output.
- Tokenization
At the start of a transformer's pipeline, input text is cut into tokens and every token is given its ID from the vocabulary; this step is tokenization.