New Benchmark Exposes Knowledge Holes in Unlearned Multimodal LLMs
A recent paper on arXiv (ID 2608.01849) tackles the issue of machine unlearning within Multimodal Large Language Models (MLLMs). The challenge lies in the fact that eliminating unsafe content can inadvertently harm the performance on benign inputs. The authors highlight a significant oversight in existing evaluation methods: they utilize benchmarks that do not adequately represent the forget set, thus missing 'knowledge holes'—notable drops in performance on benign inputs that resemble the discarded data. To investigate these knowledge gaps, the team developed a benchmark to identify such degradation and confirmed through experiments that these issues arise systematically from prevalent unlearning techniques. To address this, they introduce 'Selective Protection with Anchored Regularization,' which safeguards generic patterns through anchored activation filtering and enhances them with extra regularization.
Key facts
- The paper is available on arXiv with ID 2608.01849.
- The paper addresses machine unlearning in Multimodal Large Language Models (MLLMs).
- Current unlearning evaluation paradigms have a blind spot: they fail to capture knowledge holes.
- Knowledge holes are severe degradation on benign adjacent inputs.
- The researchers constructed a benchmark to probe knowledge holes.
- Controlled experiments confirmed knowledge holes are a systematic consequence of common unlearning approaches.
- The proposed method is called 'Selective Protection with Anchored Regularization'.
- The method protects generic patterns via anchored activation filtering.
Entities
—