Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Summary
This survey, accepted to ACL Findings 2026, provides a unified, system-oriented overview of multimodal unlearning across vision, language, audio, and video. It addresses the challenge of selectively removing sensitive, copyrighted, biased, or unsafe cross-modal associations from multimodal foundation models like VLMs, DMs, LLMs, and AFMs. Retraining these models after data deletion requests or policy updates is often impractical due to knowledge being distributed across shared representations. The survey introduces a taxonomy for systematic comparison of unlearning methods across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. It also highlights open problems and practical considerations to guide future research and deployment, and includes a curated repository.
Key takeaway
For Machine Learning Engineers deploying multimodal foundation models, understanding unlearning is crucial for compliance and safety. You should evaluate unlearning methods based on the survey's taxonomy, considering trade-offs in deletion strength, retention, and efficiency. This approach helps address sensitive, copyrighted, or biased data without expensive full model retraining. Explore the curated repository to inform your implementation strategies.
Key insights
Multimodal unlearning selectively removes unwanted data associations from foundation models without costly retraining.
Principles
- Knowledge distribution complicates targeted forgetting.
- Unlearning balances deletion strength with utility retention.
- Taxonomy clarifies trade-offs in unlearning methods.
Method
The survey proposes a taxonomy for comparing multimodal unlearning methods, evaluating trade-offs in deletion strength, retention, efficiency, reversibility, and robustness across various model architectures and modalities.
In practice
- Address sensitive data in VLMs, DMs, LLMs, AFMs.
- Evaluate unlearning methods using proposed taxonomy.
- Consult curated repository for resources.
Topics
- Multimodal Unlearning
- Foundation Models
- Data Privacy
- Model Governance
- Vision-Language Models
- Diffusion Models
- Large Language Models
Best for: AI Architect, Research Scientist, CTO, AI Scientist, Machine Learning Engineer, AI Security Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.AI updates on arXiv.org.