ARTFEED — Contemporary Art Intelligence

New MobileWorldSafety Benchmark Tests AI Agents Against Injection Attacks on Android

ai-technology · 2026-08-19

MobileWorldSafety, a new benchmark for evaluating LLM-powered GUI agents on Android devices, was introduced via an arXiv paper. The benchmark comprises 142 risk tasks developed on real Android applications, specifically targeting environmental injection attacks such as indirect prompt injections and adversarial instructions. These attacks can manipulate agent behavior unnoticed through everyday mobile scenarios. Each task includes a programmatically verifiable risk indicator based on the final system state, and outcomes are assessed through a two-stage pipeline. This initiative addresses gaps in existing benchmarks that overlook realistic user contexts. The paper, listed as arXiv:2608.17659, represents a step toward safer deployment of autonomous smartphone agents.

Key facts

  • MobileWorldSafety is a new benchmark for GUI agent safety.
  • It consists of 142 risk tasks.
  • The benchmark is built on real Android applications.
  • It focuses on environmental injection attacks.
  • Injection attacks include indirect prompt injections and adversarial instructions.
  • Each task has a programmatically verifiable risk indicator.
  • Evaluation uses a two-stage pipeline.
  • The paper is available on arXiv with identifier 2608.17659.

Entities

Institutions

  • arXiv

Sources