Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

· Source: cs.AI updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Cybersecurity & Data Privacy · Depth: Expert, quick

Summary

This survey, accepted to ACL Findings 2026, provides a unified, system-oriented overview of multimodal unlearning across vision, language, audio, and video. It addresses the challenge of selectively removing sensitive, copyrighted, biased, or unsafe cross-modal associations from multimodal foundation models like VLMs, DMs, LLMs, and AFMs. Retraining these models after data deletion requests or policy updates is often impractical due to knowledge being distributed across shared representations. The survey introduces a taxonomy for systematic comparison of unlearning methods across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. It also highlights open problems and practical considerations to guide future research and deployment, and includes a curated repository.

Key takeaway

For Machine Learning Engineers deploying multimodal foundation models, understanding unlearning is crucial for compliance and safety. You should evaluate unlearning methods based on the survey's taxonomy, considering trade-offs in deletion strength, retention, and efficiency. This approach helps address sensitive, copyrighted, or biased data without expensive full model retraining. Explore the curated repository to inform your implementation strategies.

Key insights

Multimodal unlearning selectively removes unwanted data associations from foundation models without costly retraining.

Principles

Method

The survey proposes a taxonomy for comparing multimodal unlearning methods, evaluating trade-offs in deletion strength, retention, efficiency, reversibility, and robustness across various model architectures and modalities.

In practice

Topics

Best for: AI Architect, Research Scientist, CTO, AI Scientist, Machine Learning Engineer, AI Security Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.AI updates on arXiv.org.