ARTFEED — Contemporary Art Intelligence

VLMs Tested for Geometry Clipping Detection in Game QA

other · 2026-07-29

A study on arXiv evaluates Vision-Language Models (VLMs) for detecting geometry clipping in video game quality assurance. A custom exploration agent navigates game levels to collect visual observations, and an automatic annotation pipeline provides frame-level clipping labels. Six recent VLMs—Gemini, GPT, Qwen, Gemma, Llama, and Ministral—are benchmarked under zero-shot prompting with four prompt variants. Results show VLMs can capture visual cues but produce substantial false positives on ambiguous frames like near-contact geometry and partial occlusions. Gemini-3.1-Flash achieves the best accuracy and robustness to prompt variation.

Key facts

  • Study uses VLMs for anomaly detection in game QA
  • Focus on geometry clipping detection
  • Custom exploration agent navigates game levels
  • Automatic annotation pipeline provides frame-level labels
  • Six VLMs benchmarked: Gemini, GPT, Qwen, Gemma, Llama, Ministral
  • Zero-shot prompting with four prompt variants
  • VLMs produce false positives on ambiguous frames
  • Gemini-3.1-Flash achieves best accuracy

Entities

Institutions

  • arXiv

Sources