ARTFEED — Contemporary Art Intelligence

OSReward: New Benchmark for Evaluating VLM Judges in Computer-Use Agents

ai-technology · 2026-08-07

Researchers have introduced OSReward, a benchmark designed to evaluate the reliability of vision-language models (VLMs) used as judges for computer-using agents (CUAs). The benchmark addresses a critical gap: while VLMs are increasingly employed to verify whether CUAs have completed tasks, their reliability has not been systematically studied. OSReward comprises trajectories from diverse agent backbones executing human-verified instructions across multiple platforms, labeled with ground-truth verdicts through multi-stage human annotation. The team also derived OSReward-Hard, a challenging subset. The work is detailed in a paper on arXiv (ID: 2607.28609), announced as a replacement. The study underscores the importance of standardized evaluation for cross-platform computer-use reward models, aiming to improve CUA evaluation, data curation, and reinforcement learning.

Key facts

  • OSReward is a benchmark for evaluating VLM judges on CUA trajectories.
  • Trajectories come from diverse agent backbones executing human-verified instructions across platforms.
  • Ground-truth verdicts are labeled through multi-stage human annotation.
  • OSReward-Hard is a challenge set derived from OSReward.
  • The paper is available on arXiv with ID 2607.28609.
  • The announcement type is 'replace'.
  • The study addresses the reliability of VLM judges in CUA evaluation.
  • The work aims to improve CUA evaluation, data curation, and reinforcement learning.

Entities

Institutions

  • arXiv

Sources