Home › Glossary › Vision-enabled chat model

Vision-enabled chat model

A chat model able to understand images included in the prompt, such as GPT-4o, GPT-4.1, the GPT-5 series and o-series models. An application layer cannot add this capability to a model that handles only text.

Also called large multimodal model.

Read more: Microsoft Learn

In the Ultra Transcenders books

AI-901AI-103

Each book explains Vision-enabled chat model in context, with comparison tables and the common traps.

Terms in this definition

See Vision-enabled chat model in the full glossary