ARTFEED — Contemporary Art Intelligence

VLM2Rec: Balancing Modalities in Vision-Language Models for Sequential Recommendation

ai-technology · 2026-08-13

A recent preprint on arXiv (2603.17450) presents VLM2Rec, a novel framework that leverages Vision-Language Models (VLMs) for multimodal sequential recommendations with an emphasis on collaborative filtering. The researchers highlight a critical issue: the conventional contrastive supervised fine-tuning (SFT) can exacerbate the existing modality imbalance, leading to one modality overshadowing the other, which ultimately affects recommendation precision. To counteract this, VLM2Rec encourages a more equitable use of modalities. The authors introduce a mechanism termed 'Weak-modality Penaliz' to ensure that both visual and textual inputs are effectively utilized. This work draws inspiration from the efficacy of Large Language Models (LLMs) as robust embedders and seeks to incorporate collaborative filtering signals into item representations. The paper can be accessed at https://arxiv.org/abs/2603.17450.

Key facts

  • Paper arXiv:2603.17450 introduces VLM2Rec framework.
  • VLM2Rec uses Vision-Language Models (VLMs) as CF-aware multimodal embedders for sequential recommendation.
  • Standard contrastive SFT can amplify modality imbalance, degrading one modality.
  • VLM2Rec promotes balanced modality utilization.
  • The paper proposes a 'Weak-modality Penaliz' mechanism.
  • Motivated by LLMs as high-capacity embedders.
  • Aims to fully integrate collaborative filtering signals into item representations.
  • Preprint announced as replace-cross.

Entities

Institutions

  • arXiv

Sources