FREE STUDY NOTES · AI-901

How generative AI models work: tokens, attention and next-token prediction

What happens between a prompt and a response, and how inference differs from training.

From Ultra Transcenders AI-901 by Tony Rough (publishing soon)

Before choosing a model, it helps to know what one actually does when you send it a prompt. This section walks through the ideas from training to response.

What generative AI is

Generative AI creates new content (text, images, audio, code) from a prompt. It is not supervised learning, where a model is trained on labelled examples to predict a label or a number.

Foundation models are large models pretrained with self-supervised learning: they learn from huge unlabelled collections of text by repeatedly predicting the next token. Because the “label” is just the next piece of the text itself, nobody has to label the data by hand.

Inside a transformer LLM

Most large language models (LLMs) use the transformer architecture. A prompt passes through three stages:

  1. Tokenization: the text is split into tokens (words, parts of words or punctuation).
  2. Embedding: each token becomes a multi-dimensional vector (a list of numbers). Tokens with similar meanings sit close together in this vector space.
  3. Next-token prediction: the model produces a probability distribution over possible next tokens, picks one, appends it to the text and repeats until the response is complete.

Along the way, attention weighs how much each other token in the context should influence the prediction. Figure 2.1 shows the loop, and how training differs from inference.

A prompt goes through tokenization (the text is split into tokens), then embedding (each token becomes a vector), then next-token prediction (probabilities for the possible next tokens, with attention weighing each token in the context). The model picks one token, appends it to the text and repeats until the response is complete. A lower band shows that training, done once beforehand, learns the weights, and inference uses those weights on every request without retraining the model.
Figure 2.1: A prompt passes through tokenization, embedding and next-token prediction, looping until the response is complete

Common trap: Attention is not the vector itself; the vector is the embedding. Object detection and anonymisation are not stages of an LLM either.

Training vs inference

A model is built once and then used many times, and these are two separate activities.

LLMs and SLMs

Language models come in different sizes. An LLM and a small language model (SLM) use the same kind of architecture; the difference is size, measured in number of parameters. SLMs are cheaper and faster but less broadly capable.

Limits: accuracy and training cutoff

Because a model predicts likely text rather than checking facts, its answers need care.

Get the whole book

This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · Free AI-901 glossary · All AI-901 study notes

More AI-901 study notes