VLMs Tested for Geometry Clipping Detection in Game QA
A study on arXiv evaluates Vision-Language Models (VLMs) for detecting geometry clipping in video game quality assurance. A custom exploration agent navigates game levels to collect visual observations, and an automatic annotation pipeline provides frame-level clipping labels. Six recent VLMs—Gemini, GPT, Qwen, Gemma, Llama, and Ministral—are benchmarked under zero-shot prompting with four prompt variants. Results show VLMs can capture visual cues but produce substantial false positives on ambiguous frames like near-contact geometry and partial occlusions. Gemini-3.1-Flash achieves the best accuracy and robustness to prompt variation.
Key facts
- Study uses VLMs for anomaly detection in game QA
- Focus on geometry clipping detection
- Custom exploration agent navigates game levels
- Automatic annotation pipeline provides frame-level labels
- Six VLMs benchmarked: Gemini, GPT, Qwen, Gemma, Llama, Ministral
- Zero-shot prompting with four prompt variants
- VLMs produce false positives on ambiguous frames
- Gemini-3.1-Flash achieves best accuracy
Entities
Institutions
- arXiv