CAP: New Benchmark for Cross-Site Browser Agents with Complex Actions and Perception
A new standard known as CAP has been launched to assess browser agents performing human-like web tasks across sites, necessitating intricate UI interactions and visual comprehension. This benchmark tackles two key challenges in genuine web browsing that are frequently ignored by current assessments: executing complex actions on sophisticated user interfaces and perceiving visually dynamic content, particularly in processes that involve multiple websites. CAP employs a decomposition-and-recomposition approach, initially abstracting each site into a structured site card that encapsulates user-facing features, complex execution tasks, and perceptual needs, before integrating these elements into realistic cross-site workflows. Each task is based on the actual functionalities of the websites. The related paper can be found on arXiv under the identifier 2608.08392.
Key facts
- CAP is a scalable benchmark for evaluating browser agents.
- It focuses on cross-site, human-like web tasks.
- The benchmark addresses complex actions over rich user interfaces.
- It also addresses visual perception of dynamically rendered content.
- The pipeline decomposes websites into site cards and recomposes them into workflows.
- Each task is grounded in the actual website's functionality.
- The paper is available on arXiv with identifier 2608.08392.
- The benchmark is designed to evaluate agents in workflows spanning multiple websites.
Entities
Institutions
- arXiv