Multimodal AI
Models that understand and generate text, image, audio and video.
Beginner5 min read
Multimodal AI
A multimodal model is an AI system that natively understands and/or generates multiple modalities such as text, image, audio and video.
Read definition →
Beginner5 min read
Vision Model
A vision model is a neural network trained to process and interpret visual data, such as photographs, screenshots or diagrams, often paired with a language model so it can describe or reason about that visual content in text.
Read definition →
Other categories