Smaller than an LLM, with roughly under 10 billion parameters but a similar kind of architecture. Running it costs less and is quicker, at the price of narrower capability.
Also called small language model.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains SLM in context, with comparison tables and the common traps.
Terms in this definition
- LLM
Large language model, usually a transformer network with billions of parameters that learnt to predict the next token from vast amounts of text. Compared with a small language model it is more capable across tasks, but slower and costlier.
- Capability
Something that a person, organisation or system is able to do.
Related terms
- GPT-4 Turbo
GPT-4 chat model, since retired in Azure, that was older and pricier than its successors. Narrow text tasks are better served, at lower cost, by a small language model.
- Model parameters
The weights a model learns in training, which together hold what it knows. How many there are is what distinguishes a small language model from a large one.
- Phi-3-mini
Small language model from Microsoft supporting a context of up to 128k tokens, with a 4k version as well. It has been retired and Phi-4-mini-instruct succeeds it.