ARTFEED — Contemporary Art Intelligence

SafeCap: Enhancing LVLM Safety via Image Captioning Reinforcement Learning

ai-technology · 2026-08-13

A recent study presents SafeCap, a reinforcement-learning framework aimed at enhancing the safety of large vision-language models (LVLMs) against jailbreak threats. This framework trains a policy model to initially create a safety-focused image caption, which is then used to derive a final answer, ensuring that the caption helps a frozen LLM arrive at a safety-oriented conclusion. This method prompts the model to highlight visual indicators pertinent to generating safe responses instead of depending solely on direct refusal guidance. SafeCap shows significant improvements in overall safety performance under its DirectCap protocol across five multimodal safety benchmarks and six vision-utility benchmarks, achieving safety average enhancements of 3.7-19.0 points across four model configurations while maintaining similar or better utility. The paper can be found on arXiv under identifier 2608.10513.

Key facts

  • SafeCap is a reinforcement-learning framework for LVLMs.
  • It trains a policy model to generate a safety-relevant image caption before producing a final answer.
  • The caption is optimized to enable a frozen LLM to reach a safety-aligned decision.
  • SafeCap improves safety performance by 3.7-19.0 points across four model settings.
  • The framework was evaluated on five multimodal safety benchmarks and six vision-utility benchmarks.
  • The paper is available on arXiv with identifier 2608.10513.
  • The approach focuses on exposing visual cues relevant to safe response generation.
  • SafeCap maintains comparable or improved utility while enhancing safety.

Entities

Institutions

  • arXiv

Sources