MobileJudgeBench: Benchmarking LLM Judges for Mobile Agent Evaluation
A novel benchmark called MobileJudgeBench has been launched to assess the effectiveness of LLM-based judges in evaluating mobile agent trajectories. This benchmark includes 931 trajectories annotated by humans, covering 6 mobile agent benchmarks, 4 agent models, and 68 applications. The research examines 6 judging techniques, featuring five modified from SPA-Bench, A3 with two configurations, AndroidArena, and AgentRewardBench, along with a straightforward baseline, across various LLM backends. Notably, findings indicate that the basic baseline judge using sampled screenshots is on par with, and frequently surpasses, specialized methods, suggesting that more complex judging systems do not necessarily enhance quality. The LLM backbone emerges as the key factor among competitive methods. The study also discusses benchmark quality, with the abstract truncated. The paper can be found on arXiv under identifier 2608.11434.
Key facts
- MobileJudgeBench is a new benchmark for evaluating LLM-as-judge methods on mobile agent trajectories.
- It comprises 931 human-annotated trajectories.
- The trajectories span 6 mobile agent benchmarks, 4 agent models, and 68 apps.
- 6 judge methods are evaluated, including five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline.
- A simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods.
- More elaborate judge pipelines do not consistently improve judge quality.
- Among competitive methods, the LLM backbone is the primary driver of judge quality.
- The paper is available on arXiv under the identifier 2608.11434.
Entities
Institutions
- arXiv