ARTFEED — Contemporary Art Intelligence

RC-GRPO: New RL Strategy Teaches MLLMs to Refuse Nonexistent Objects

ai-technology · 2026-08-06

A new reinforcement learning approach called Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) has been developed by researchers to enhance the refusal capabilities of Multimodal Large Language Models (MLLMs) in Generalized Referring Expression Comprehension (GREC) tasks. GREC involves identifying objects based on textual descriptions, requiring models to recognize both existing (positive samples) and nonexistent (negative samples) items. Although MLLMs perform well in locating real objects, they struggle to reject those that aren't present due to a lack of negative samples in training, which results in inaccurate bounding boxes. Current post-training methods like supervised fine-tuning (SFT) and reinforcement learning (RL) improve refusal skills but often compromise localization accuracy. RC-GRPO resolves this by fine-tuning the RL approach to enhance refusal without sacrificing localization performance. This research, which is crucial for the reliability of MLLMs in visual grounding tasks, is detailed in a paper on arXiv (arXiv:2608.04698) and is noted as a cross-type submission.

Key facts

  • RC-GRPO is a new reinforcement learning strategy for MLLMs.
  • It targets Generalized Referring Expression Comprehension (GREC).
  • GREC requires models to localize objects when present and refuse when absent.
  • MLLMs often hallucinate bounding boxes for nonexistent objects.
  • Existing SFT and RL methods improve refusal but degrade localization accuracy.
  • RC-GRPO calibrates RL to preserve localization while enhancing refusal.
  • The paper is available on arXiv with ID 2608.04698.
  • The announcement type is 'cross'.

Entities

Institutions

  • arXiv

Sources