Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cybersecurity & Data Privacy · Depth: Expert, quick

Summary

A new survey provides a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video. This field addresses the challenge of foundation models like VLMs, DMs, LLMs, and AFMs inadvertently encoding sensitive, copyrighted, biased, or unsafe cross-modal associations from their training data. Retraining these models after deletion requests or policy updates is often impractical, and targeted forgetting is difficult due to knowledge distribution across shared representations. Multimodal unlearning enables selective removal of specific information across modalities while retaining overall model utility. The survey establishes a taxonomy for systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. It also highlights open problems and practical considerations to guide future research and deployment. A curated repository is available at https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/.

Key takeaway

For AI Scientists and Machine Learning Engineers developing or deploying multimodal foundation models, you must consider robust multimodal unlearning strategies. This is crucial for addressing inadvertent encoding of sensitive or biased data without costly full model retraining. You should evaluate unlearning methods based on the survey's taxonomy, prioritizing deletion strength, efficiency, and retention to ensure compliance and model integrity. Explore the curated repository to inform your implementation decisions.

Key insights

Multimodal unlearning selectively removes sensitive data from foundation models while preserving utility.

Principles

Method

The survey proposes a taxonomy for comparing multimodal unlearning methods based on model architectures and modalities, evaluating deletion strength, retention, efficiency, reversibility, and robustness.

In practice

Topics

Best for: Research Scientist, CTO, VP of Engineering/Data, AI Scientist, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.