Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Summary
A new survey provides a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video. This field addresses the challenge of foundation models like VLMs, DMs, LLMs, and AFMs inadvertently encoding sensitive, copyrighted, biased, or unsafe cross-modal associations from their training data. Retraining these models after deletion requests or policy updates is often impractical, and targeted forgetting is difficult due to knowledge distribution across shared representations. Multimodal unlearning enables selective removal of specific information across modalities while retaining overall model utility. The survey establishes a taxonomy for systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. It also highlights open problems and practical considerations to guide future research and deployment. A curated repository is available at https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/.
Key takeaway
For AI Scientists and Machine Learning Engineers developing or deploying multimodal foundation models, you must consider robust multimodal unlearning strategies. This is crucial for addressing inadvertent encoding of sensitive or biased data without costly full model retraining. You should evaluate unlearning methods based on the survey's taxonomy, prioritizing deletion strength, efficiency, and retention to ensure compliance and model integrity. Explore the curated repository to inform your implementation decisions.
Key insights
Multimodal unlearning selectively removes sensitive data from foundation models while preserving utility.
Principles
- Knowledge in foundation models is distributed across shared representations.
- Retraining for data deletion is often impractical.
- Unlearning involves trade-offs: strength, retention, efficiency, reversibility, robustness.
Method
The survey proposes a taxonomy for comparing multimodal unlearning methods based on model architectures and modalities, evaluating deletion strength, retention, efficiency, reversibility, and robustness.
In practice
- Address sensitive data encoding in VLMs, DMs, LLMs, AFMs.
- Implement selective data removal for policy compliance.
- Evaluate unlearning methods using the proposed taxonomy.
Topics
- Multimodal Unlearning
- Foundation Models
- Data Privacy
- Model Debiasing
- Machine Unlearning
- VLMs, LLMs, DMs, AFMs
Best for: Research Scientist, CTO, VP of Engineering/Data, AI Scientist, Machine Learning Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.