For Heads-of · Practitioner
Multimodal AI
AI systems that process and integrate multiple types of data (text, images, audio, video) to understand and respond to the world.
- multimodal
- vision
- audio
- integration
Capabilities
Multimodal models can answer questions about images, describe videos, generate images from text, or translate between modalities in ways that broader than single-modality models.
Complexity
Aligning different modalities and understanding their interactions increases model complexity and makes behavior harder to predict or explain.
Governance
Multimodal systems may inherit risks from all input modalities. For example, an image-text model combines visual bias and linguistic bias, potentially amplifying fairness risks.