OSReward: New Benchmark for Evaluating VLM Judges in Computer-Use Agents
Researchers have introduced OSReward, a benchmark designed to evaluate the reliability of vision-language models (VLMs) used as judges for computer-using agents (CUAs). The benchmark addresses a critical gap: while VLMs are increasingly employed to verify whether CUAs have completed tasks, their reliability has not been systematically studied. OSReward comprises trajectories from diverse agent backbones executing human-verified instructions across multiple platforms, labeled with ground-truth verdicts through multi-stage human annotation. The team also derived OSReward-Hard, a challenging subset. The work is detailed in a paper on arXiv (ID: 2607.28609), announced as a replacement. The study underscores the importance of standardized evaluation for cross-platform computer-use reward models, aiming to improve CUA evaluation, data curation, and reinforcement learning.
Key facts
- OSReward is a benchmark for evaluating VLM judges on CUA trajectories.
- Trajectories come from diverse agent backbones executing human-verified instructions across platforms.
- Ground-truth verdicts are labeled through multi-stage human annotation.
- OSReward-Hard is a challenge set derived from OSReward.
- The paper is available on arXiv with ID 2607.28609.
- The announcement type is 'replace'.
- The study addresses the reliability of VLM judges in CUA evaluation.
- The work aims to improve CUA evaluation, data curation, and reinforcement learning.
Entities
Institutions
- arXiv