OpenAI chat model that is multimodal, taking images and text and, in some variants, audio. Temperature and penalty parameters work with it, and the 2024-08-06 version allows supervised, DPO and vision fine-tuning.
Also called GPT-4o.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains GPT-4o in context, with comparison tables and the common traps.
Terms in this definition
- Model temperature
Sampling setting between 0 and 2 controlling how random the output is. Lower values yield focused, repeatable responses and higher ones more creative text; tune this or Top P, but not both.
- DPO
Direct preference optimization, a fine-tuning technique that trains on pairs of better and worse responses without needing a reward model. Records contain input, preferred_output and non_preferred_output, and the two outputs can't be identical.
- Vision fine-tuning
Fine-tuning where the JSONL messages include images as URLs or base64. Only gpt-4o (2024-08-06) and gpt-4.1 (2025-04-14) support it, and images showing people, faces or CAPTCHAs are dropped.
Related terms
- GPT-4 vision-preview
Preview model adding vision to GPT-4, with OCR, object grounding and video enhancements. Now retired, it gave way to gpt-4 turbo-2024-04-09 and later GPT-4o, so it should not be chosen for new image scenarios.
- Vision-enabled chat model
A chat model able to understand images included in the prompt, such as GPT-4o, GPT-4.1, the GPT-5 series and o-series models. An application layer cannot add this capability to a model that handles only text.