For Heads-of · Practitioner

Multimodal AI

AI systems that process and integrate multiple types of data (text, images, audio, video) to understand and respond to the world.

  • multimodal
  • vision
  • audio
  • integration

Capabilities

Multimodal models can answer questions about images, describe videos, generate images from text, or translate between modalities in ways that broader than single-modality models.

Complexity

Aligning different modalities and understanding their interactions increases model complexity and makes behavior harder to predict or explain.

Governance

Multimodal systems may inherit risks from all input modalities. For example, an image-text model combines visual bias and linguistic bias, potentially amplifying fairness risks.

← Back to Glossary