A generative model able to take several kinds of input, for example images alongside text, within one request.
Also called vision-enabled model.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Multimodal model in context, with comparison tables and the common traps.
Related terms
- Alt text (Image Analysis)
Image description for accessibility, produced by Image Analysis captioning above a confidence threshold. New solutions should prefer Content Understanding or a multimodal model.