ARTFEED — Contemporary Art Intelligence

CAP: New Benchmark for Cross-Site Browser Agents with Complex Actions and Perception

ai-technology · 2026-08-11

A new standard known as CAP has been launched to assess browser agents performing human-like web tasks across sites, necessitating intricate UI interactions and visual comprehension. This benchmark tackles two key challenges in genuine web browsing that are frequently ignored by current assessments: executing complex actions on sophisticated user interfaces and perceiving visually dynamic content, particularly in processes that involve multiple websites. CAP employs a decomposition-and-recomposition approach, initially abstracting each site into a structured site card that encapsulates user-facing features, complex execution tasks, and perceptual needs, before integrating these elements into realistic cross-site workflows. Each task is based on the actual functionalities of the websites. The related paper can be found on arXiv under the identifier 2608.08392.

Key facts

  • CAP is a scalable benchmark for evaluating browser agents.
  • It focuses on cross-site, human-like web tasks.
  • The benchmark addresses complex actions over rich user interfaces.
  • It also addresses visual perception of dynamically rendered content.
  • The pipeline decomposes websites into site cards and recomposes them into workflows.
  • Each task is grounded in the actual website's functionality.
  • The paper is available on arXiv with identifier 2608.08392.
  • The benchmark is designed to evaluate agents in workflows spanning multiple websites.

Entities

Institutions

  • arXiv

Sources