Vision-Language Models Overestimate Contextual Cues in Cross-Market Ad Preference Prediction
A recent paper on arXiv (2608.04504) reveals a failure mode in vision-language models (VLMs) termed Contextual Variable Overestimation (CVE). This phenomenon occurs when models prioritize prominent visual and textual signals while neglecting crucial but infrequent contextual variables. This problem is particularly noticeable in assessing preferences for advertisement images across various geographic regions. For instance, when tasked with selecting between two product images designed for different countries, a VLM tends to produce a uniform response, disregarding actual regional differences. This failure arises as dominant signals, like product features and dense image segments, overshadow the few essential tokens that represent specific market contexts. To combat CVE, researchers compiled a new multimodal dataset featuring real advertising creatives and their click-through rates. The paper has been published on arXiv and is noted as a cross-type submission.
Key facts
- Paper arXiv:2608.04504 identifies Contextual Variable Overestimation (CVE) in vision-language models.
- CVE causes VLMs to overestimate dominant cues and underestimate sparse contextual variables.
- The issue is observed in cross-market advertisement image preference prediction.
- VLMs often default to consistent outputs when comparing product images for different countries.
- High-volume signals like product attributes and dense image patches overwhelm critical market-specific tokens.
- The researchers collected a new multimodal dataset of real advertising creatives and click-through performance.
- The paper is announced as a cross-type submission on arXiv.
Entities
Institutions
- arXiv