ARTFEED — Contemporary Art Intelligence

UNLINK-VL: New Benchmark for Cross-Modal Knowledge Unlearning in Vision-Language Models

ai-technology · 2026-08-06

A new benchmark called UNLINK-VL has been developed by researchers to assess cross-modal knowledge unlearning in Vision-Language Models (VLMs). This benchmark fills a void in current unlearning research, which has mainly concentrated on forgetting within single modalities. Operating in a post-hoc unlearning framework, UNLINK-VL lacks access to original forget and retain corpora and instead targets visually identifiable real-world entities. These entities are linked to relevant images and one-hop and multi-hop facts sourced from Wikidata. The benchmark includes four distinct subsets that analyze various facets of cross-modal unlearning. This research is crucial for creating reliable AI systems by eliminating sensitive, copyrighted, or harmful information from VLMs, akin to Large Language Models (LLMs). The paper can be found on arXiv with the identifier 2608.03791.

Key facts

  • UNLINK-VL is a real-world benchmark for cross-modal knowledge unlearning in Vision-Language Models (VLMs).
  • It addresses the gap in existing unlearning studies that focus on forgetting within individual modalities.
  • The benchmark operates under a post-hoc unlearning setting where original forget and retain corpora are unavailable.
  • It selects visually identifiable real-world entities as unlearning targets.
  • Entities are associated with corresponding images and one-hop and multi-hop facts derived from Wikidata.
  • The benchmark comprises four complementary subsets for evaluation.
  • The work aims to remove sensitive, copyrighted, or harmful knowledge from VLMs.
  • The paper is available on arXiv with identifier 2608.03791.

Entities

Institutions

  • arXiv

Sources