Models that understand and generate text, image, audio and video.
A multimodal model is an AI system that natively understands and/or generates multiple modalities such as text, image, audio and video.