SPIEval: Benchmarking LLMs as Mobile Assistants on Scattered Personal Data
A new evaluation standard, SPIEval, has been launched to assess large language models (LLMs) functioning as mobile assistants that manage personal data across various applications. This benchmark, outlined in a paper on arXiv (2608.10692), is curated by humans and focuses on five cognitive skills: reasoning, disambiguation, integration, preference inference, and multi-intent decomposition. It consists of 250 tasks that utilize 4,335 personal records from 10 different apps, facilitating multi-turn interactions through 21 tools. The findings highlight a range of scenarios, challenging tasks, fragmented information, controllable settings, and measurable results. Among nine evaluated LLMs, the highest scorer, GPT-5.5 (xhigh), received a modest rating, suggesting significant potential for enhancement. This benchmark seeks to address the need for specialized evaluation of LLMs as mobile assistants, offering a consistent method to gauge their capability in utilizing dispersed personal information.
Key facts
- SPIEval is a human-curated benchmark for evaluating LLMs as mobile assistants.
- It focuses on leveraging personal information scattered across multiple apps.
- The benchmark is grounded in five cognitive capabilities: reasoning, disambiguation, integration, preference inference, and multi-intent decomposition.
- SPIEval includes 250 tasks, 4,335 personal records, 10 apps, and 21 tools for multi-turn interaction.
- The benchmark supports multi-turn interaction through 21 tools.
- Nine representative LLMs were evaluated.
- The best-performing model was GPT-5.5 (xhigh), but it still showed substantial room for improvement.
- The paper is available on arXiv with identifier 2608.10692.
Entities
—