ParaGUIBench: Benchmarking Parallel GUI Agent Coordination
Researchers have introduced ParaGUIBench, the first benchmark for parallel execution and coordination of multiple GUI agents. These agents, powered by large multimodal models (LMMs), perform actions like clicking and typing on desktops and mobile devices. Current agents struggle with long-horizon tasks due to costly LMM inferences and performance degradation. ParaGUIBench enables parallel coordination among agents on separate desktop instances, inspired by human collaboration. It includes a multi-device Docker infrastructure with a shared file system, a dataset of 233 tasks across six categories, and an evaluation system with efficiency metrics like step reduction. This work addresses a gap in GUI agent research, aiming to improve scalability and performance.
Key facts
- ParaGUIBench is the first benchmark for parallel execution and coordination of multiple GUI agents.
- GUI agents are powered by large multimodal models (LMMs).
- Current agents scale poorly to long-horizon tasks due to costly LMM inferences and context growth.
- The benchmark includes a multi-device Docker infrastructure with a shared file system.
- The dataset contains 233 tasks spanning six task categories.
- The evaluation system includes efficiency metrics such as step reduction.
- Parallel coordination among GUI agents has received little attention prior to this work.
- The approach is inspired by human division of workloads among collaborators.
Entities
Institutions
- arXiv