ARTFEED — Contemporary Art Intelligence

Look Twice: Training-Free Evidence Highlighting for KB-VQA

ai-technology · 2026-08-07

A recent study presents Look Twice (LoT), an inference-time framework that enhances Knowledge-based Visual Question Answering (KB-VQA) without the need for training. This approach utilizes the internal attention of the model to emphasize pertinent evidence, tackling issues related to noisy retrieval and distracting visual elements that can lead Multimodal Large Language Models (MLLMs) to miss crucial information. LoT effectively pinpoints image regions and text sentences relevant to the query, eliminating attention sinks and distractions, while reformulating the input to underscore the chosen evidence prior to generating answers. This framework operates without requiring parameter adjustments, additional models, or changes to architecture, making it a straightforward plug-and-play option. The paper is accessible on arXiv with identifier 2604.01280 and is categorized as a replace-cross type, contributing to AI, machine learning, and multimodal understanding, particularly in art-related visual question answering systems.

Key facts

  • Look Twice (LoT) is a training-free inference-time framework for KB-VQA.
  • It uses the model's internal attention to select evidence.
  • It filters attention sinks and distracting content.
  • It reformulates input to highlight selected evidence.
  • No parameter updates, auxiliary models, or architectural modifications are required.
  • The paper is available on arXiv with ID 2604.01280.
  • The announcement type is replace-cross.
  • The method aims to improve MLLMs' performance on KB-VQA tasks.

Entities

Institutions

  • arXiv

Sources