StepJack: New Benchmark Exposes Multi-Step Injection Risks in Computer-Use Agents
Researchers have introduced StepJack, a safety benchmark designed to evaluate computer-use agents (CUAs) against a new class of attack called multi-step indirect prompt injection. This attack decomposes a malicious goal into multiple innocuous-looking sub-steps distributed across a chain of web pages referenced along the agent's navigation path. The team developed an automated pipeline that decomposes adversarial goals while ensuring the sub-steps achieve the original goal and appear harmless. StepJack comprises 480 test examples. Evaluation of six state-of-the-art CUAs revealed that multi-step attacks increase the attack success rate (ASR) on three of the six agents, by up to 31.2 percentage points at a fixed decomposition depth. The findings highlight a significant vulnerability in current CUAs, which are increasingly used for tasks like web browsing and form filling. The benchmark aims to spur development of more robust defenses. The paper is available on arXiv under identifier 2608.06477.
Key facts
- StepJack is a new safety benchmark for computer-use agents (CUAs).
- It targets multi-step indirect prompt injection attacks.
- The attack decomposes adversarial goals into innocuous sub-steps.
- Sub-steps are distributed across a chain of web pages.
- The pipeline automatically decomposes adversarial goals.
- StepJack includes 480 test examples.
- Six state-of-the-art CUAs were evaluated.
- Multi-step attacks raised ASR on three of six CUAs, by up to 31.2 points.
- The paper is on arXiv with ID 2608.06477.
Entities
Institutions
- arXiv