ARTFEED — Contemporary Art Intelligence

GAM-Agent: Game-Theoretic Multi-Agent Framework Enhances Vision-Language Reasoning

ai-technology · 2026-08-13

A new multi-agent framework named GAM-Agent has been developed by researchers to enhance vision-language reasoning. Unlike earlier models that rely on a single agent, GAM-Agent approaches reasoning as a non-zero-sum game, involving base agents that focus on specific visual perception tasks and a key agent responsible for ensuring logical consistency and factual accuracy. These agents exchange structured claims, evidence, and uncertainty assessments. An adaptive controller responds to disagreements or ambiguities by initiating multi-round discussions, resulting in more reliable and interpretable outcomes. Testing on four demanding benchmarks—MMMU, MMBench, MVBench, and V*Bench—shows that GAM-Agent significantly outperforms various vision-language model (VLM) backbones, particularly improving the accuracy of smaller models like Qwen2.5-VL-7B and InternVL2.5-8B, achieving results on par with or surpassing larger counterparts. This game-theoretic framework presents a fresh strategy for advancing AI reasoning, with promising implications for intricate visual reasoning challenges.

Key facts

  • GAM-Agent is a game-theoretic multi-agent framework for vision-language reasoning.
  • It models reasoning as a non-zero-sum game between base agents and a critical agent.
  • Agents communicate via structured claims, evidence, and uncertainty estimates.
  • An uncertainty-aware controller dynamically adjusts collaboration and triggers debates.
  • Tested on MMMU, MMBench, MVBench, and V*Bench benchmarks.
  • Improves performance across various VLM backbones.
  • Boosts accuracy of small-to-mid scale models like Qwen2.5-VL-7B and InternVL2.5-8B.
  • Published on arXiv with ID 2505.23399.

Entities

Institutions

  • arXiv

Sources