Home › Glossary › Multimodal model

Multimodal model

A generative model able to take several kinds of input, for example images alongside text, within one request.

Also called vision-enabled model.

Read more: Microsoft Learn

In the Ultra Transcenders books

AI-901AI-103

Each book explains Multimodal model in context, with comparison tables and the common traps.

Related terms

See Multimodal model in the full glossary