ARTFEED — Contemporary Art Intelligence

Cross-Modal Visual Feedback Improves Prompt Optimization for VLMs

ai-technology · 2026-07-29

A recent study presents Cross-Modal Visual Feedback (CMVF) to tackle a significant drawback in automatic prompt optimization (APO) for vision-language models (VLMs). Existing APO techniques utilize a blind feedback mechanism, where the optimizer processes the question, prediction, and correct answer but does not examine the input image, which limits the ability to identify visually grounded mistakes. CMVF introduces a visual diagnosis phase conditioned on failure, allowing a more robust optimizer VLM to analyze each unsuccessful image without seeing predictions or labels. This is followed by an error-aware aggregation phase that distills findings into reusable visual blind-spot patterns for prompt revision. The image is only utilized during optimization, while the final output remains a standard prompt. The paper can be found on arXiv with ID 2607.24354.

Key facts

  • Paper ID: arXiv:2607.24354
  • Published on arXiv
  • Introduces Cross-Modal Visual Feedback (CMVF)
  • Addresses blind feedback channel in automatic prompt optimization
  • CMVF includes failure-conditioned visual diagnosis stage
  • CMVF includes error-aware aggregation stage
  • Image consumed only during optimization
  • Deployed artifact is an ordinary prompt

Entities

Institutions

  • arXiv

Sources