Large language model, usually a transformer network with billions of parameters that learnt to predict the next token from vast amounts of text. Compared with a small language model it is more capable across tasks, but slower and costlier.
Also called large language model.
Read more: Microsoft Learn
In the Ultra Transcenders books
AI-901AI-103AI-200DP-900AB-900SC-200
Each book explains LLM in context, with comparison tables and the common traps.
Terms in this definition
- Transformer
The neural network design underlying GPT and most other LLMs. Text is tokenised and embedded, attention is applied across the context, and the next token is predicted.
- PREDICT
A Fabric function for scoring MLflow models in batches from a notebook, provided the models have signatures; you can call it via MLFlowTransformer, Spark SQL or a PySpark UDF. Fabric Warehouse does not support the T-SQL PREDICT statement.
- NEXT
Restricted to visual calculations, this DAX function reads the value one step further along an axis of the visual matrix; writing OFFSET with 1 gives the same result.
- Token
The unit of text an LLM works with, which may be a word, part of a word or punctuation. Billing, limits and context windows are all counted in these units.
- SLM
Smaller than an LLM, with roughly under 10 billion parameters but a similar kind of architecture. Running it costs less and is quicker, at the price of narrower capability.
Related terms
- Agentic retrieval
Query approach in Azure AI Search where an LLM turns a conversation into several planned subqueries; these execute together, are reranked semantically, and the best chunks are merged. Classic RAG, by contrast, issues just one query.
- Fabrication
Convincing yet wrong content produced by an LLM, which does not check facts. Grounding lowers how often it happens but can't remove it entirely.
- Prompt flow
Tool in Foundry and Azure Machine Learning for assembling LLM applications as graphs of nodes (LLM, prompt and Python tools), with variants and evaluation. Microsoft Agent Framework is replacing it as it is retired.
- Prompty
A file format, open and from Microsoft, that captures an LLM prompt as a template beneath front matter specifying the model and its parameters, so prompts can live under version control in Git.
- Semantic caching
A pattern, in an application or gateway, that returns a previous LLM answer when a new prompt means much the same thing, using RediSearch vector search. Azure Managed Redis doesn't offer it as a feature in its own right.
- Span
One operation inside a trace, such as a single LLM call or prompt flow node, carrying attributes plus start and end times. Nested spans reveal the order of calls.