ARTFEED — Contemporary Art Intelligence

Test-Time Self-Evolving Framework Enhances GUI Visual Grounding

other · 2026-08-13

A new research paper on arXiv (ID: 2608.11191) introduces a Test-Time Self-Evolving framework for GUI Visual Grounding, a fundamental capability for GUI agents. The framework enables models to improve after deployment without human-annotated ground truth, addressing the limitation of existing models that freeze parameters post-deployment and cannot adapt to unseen interfaces. The proposed method constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. During exploration, the agent predicts grounding coordinates for given instructions on unseen interfaces. An MLLM-based Reflector evaluates these predictions and provides reasoning reflections. To internalize this reflection knowledge into model weights, the paper proposes Reflection-Guided On-Policy Self-Distillation. The approach overcomes the inability of recent test-time reinforcement learning methods to reflect upon failed exploration. The paper is categorized as a cross-type announcement and is available on arXiv.

Key facts

  • Paper ID: arXiv:2608.11191
  • Title: Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
  • Proposes a Test-Time Self-Evolving framework for GUI Visual Grounding
  • Framework includes Exploration, Evaluation, Reflection, and Internalization
  • Uses an MLLM-based Reflector for evaluation and reflection
  • Introduces Reflection-Guided On-Policy Self-Distillation for internalization
  • Aims to adapt models to unseen interfaces without human-annotated ground truth
  • Addresses limitations of existing test-time reinforcement learning methods

Entities

Institutions

  • arXiv

Sources