A chat model able to understand images included in the prompt, such as GPT-4o, GPT-4.1, the GPT-5 series and o-series models. An application layer cannot add this capability to a model that handles only text.
Also called large multimodal model.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Vision-enabled chat model in context, with comparison tables and the common traps.
Terms in this definition
- GPT-4o
OpenAI chat model that is multimodal, taking images and text and, in some variants, audio. Temperature and penalty parameters work with it, and the 2024-08-06 version allows supervised, DPO and vision fine-tuning.
- GPT-4.1
OpenAI chat models gpt-4.1, gpt-4.1-mini and gpt-4.1-nano, which accept images and text and return text. They are deprecated in Azure: nano retires on 14 October 2026, while the full and mini models follow on 14 April 2027.
- GPT-5 series
A series of OpenAI reasoning models, including gpt-5, mini and nano and their successors, that take text and images. Penalty and temperature parameters are not supported.
- o-series
Reasoning models from OpenAI, namely o1, o3 and o4-mini, that think longer to tackle science, coding and maths problems. Azure has deprecated all of them, with retirement on 19 November 2026 and GPT-5.6 models as successors.
- Architecture Definition Document
A key deliverable bringing together the main architecture artifacts across the four domains for every relevant state: baseline, transition and target. It sets out, in qualitative terms, what the architect intends.
- Capability
Something that a person, organisation or system is able to do.