Core
Multimodal AI
AI models capable of processing and generating multiple types of data — text, images, audio, and video — within a single system.
Explained at five levels
Level 1
AI that can understand pictures, text, and sounds all at once — not just reading, but also seeing and hearing.
Level 2
AI that works with more than just text — it can also understand images, audio, and video. Like how you can both read and look at photos.
Level 3
AI models capable of processing and generating multiple types of data — text, images, audio, and video — within a single system.
Level 4
Models that accept and produce content across modalities (text, images, audio, video) through unified architectures, enabling cross-modal reasoning and generation.
Level 5
Architectures that learn joint representations across heterogeneous data modalities via shared latent spaces or cross-attention fusion, enabling zero-shot cross-modal transfer and compositional multimodal reasoning.
Definitions are educational summaries. Terminology can vary by source and context.
Sources
- NIST Trustworthy and Responsible AI Resource Center: Glossary — National Institute of Standards and Technology. Accessed 2026-07-20.