Fine-tuning where the JSONL messages include images as URLs or base64. Only gpt-4o (2024-08-06) and gpt-4.1 (2025-04-14) support it, and images showing people, faces or CAPTCHAs are dropped.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Vision fine-tuning in context, with comparison tables and the common traps.
Terms in this definition
- Fine-tuning
Continuing to train a pretrained base model on task examples, through supervised fine-tuning, DPO or RFT, so its weights shift towards a style, format or task. It neither acts as a safety control nor adds new knowledge well.
- WHERE
Limits a SELECT, UPDATE or DELETE to just the rows meeting a condition. Omit it, and the statement hits every row.
- JSONL
Format in which each line holds a single JSON object. Fine-tuning training data and batch input files must use it, so CSV, TSV or one big JSON array may be rejected; ordinary REST calls don't involve it.
- Agents (classic) API
First-generation Foundry Agent Service API, based on threads, messages and runs. It is deprecated, replaced by conversations and responses, and retires on 31 March 2027.
- GPT-4o
OpenAI chat model that is multimodal, taking images and text and, in some variants, audio. Temperature and penalty parameters work with it, and the 2024-08-06 version allows supervised, DPO and vision fine-tuning.
- GPT-4.1
OpenAI chat models gpt-4.1, gpt-4.1-mini and gpt-4.1-nano, which accept images and text and return text. They are deprecated in Azure: nano retires on 14 October 2026, while the full and mini models follow on 14 April 2027.
Related terms
- GPT-3.5 Turbo
An older OpenAI chat model that takes text in and returns text. Because it cannot read images or transcribe audio, it is unsuitable for vision work or vision fine-tuning.