How many tokens, counting input and output together, a model can handle in one request; GPT-4.1-mini manages roughly 1M, for instance, against 128k for the now-retired Phi-3-mini.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Context window in context, with comparison tables and the common traps.
Terms in this definition
- GPT-4.1-mini
An Azure OpenAI chat model that is quicker and cheaper, accepts images as well as text and offers a context window of roughly 1M tokens, smaller on certain deployment types. It can be fine-tuned with SFT or DPO.
- Phi-3-mini
Small language model from Microsoft supporting a context of up to 128k tokens, with a 4k version as well. It has been retired and Phi-4-mini-instruct succeeds it.
Related terms
- Response compaction
Operation in the Responses API that condenses the context of a lengthy conversation into compaction items, leaving room in the context window for further turns.