What happens between a prompt and a response, and how inference differs from training.
From Ultra Transcenders AI-901 by Tony Rough (publishing soon)
Before choosing a model, it helps to know what one actually does when you send it a prompt. This section walks through the ideas from training to response.
Generative AI creates new content (text, images, audio, code) from a prompt. It is not supervised learning, where a model is trained on labelled examples to predict a label or a number.
Foundation models are large models pretrained with self-supervised learning: they learn from huge unlabelled collections of text by repeatedly predicting the next token. Because the “label” is just the next piece of the text itself, nobody has to label the data by hand.
Most large language models (LLMs) use the transformer architecture. A prompt passes through three stages:
Along the way, attention weighs how much each other token in the context should influence the prediction. Figure 2.1 shows the loop, and how training differs from inference.
Common trap: Attention is not the vector itself; the vector is the embedding. Object detection and anonymisation are not stages of an LLM either.
A model is built once and then used many times, and these are two separate activities.
Language models come in different sizes. An LLM and a small language model (SLM) use the same kind of architecture; the difference is size, measured in number of parameters. SLMs are cheaper and faster but less broadly capable.
Because a model predicts likely text rather than checking facts, its answers need care.
This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-901 glossary · All AI-901 study notes
Fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability, and how to tell them apart in a scenario.
How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.
Which model setting controls randomness, which controls length and cost, and which ones are not set at deployment.
How to recognise each AI workload from a scenario.
Agents as model plus instructions, knowledge and tools, and the auto, required and none tool_choice values.
Speech to text, text to speech, translation, batch transcription and speaker recognition compared.
How analyzers turn documents, images, audio and video into structured JSON, and how to call them from code.