ARTFEED — Contemporary Art Intelligence

UserToolBench: Benchmarking Personalized Tool-Use in LLMs

ai-technology · 2026-08-13

UserToolBench, detailed in arXiv:2608.10042, serves as a novel benchmark aimed at assessing personalized decision-making capabilities in tool-use large language models (LLMs). Unlike prior benchmarks that emphasize profile recall, style mimicry, or general tool application, UserToolBench evaluates a model's ability to deduce implicit user preferences from interaction history, identify when clarification is necessary, and generate user-aligned tool-call sequences despite incomplete data. This benchmark is constructed from privacy-protected real interaction data and integrates structured persona profiles, public API-style tool ecosystems, and extended multi-turn trajectories. It encompasses 10 user profiles, 36 tool sets, 1,065 turns, and 170 distinct tools, with tasks addressing lack-of-information, single-tool, and multi-tool scenarios. Tests on advanced tool-use LLMs indicate that these models still face challenges in executing personalized decision-making tasks effectively.

Key facts

  • UserToolBench is a benchmark for personalized decision making in tool-use LLMs.
  • It tests inference of latent user preferences from interaction history.
  • It includes recognition of when clarification is needed.
  • It evaluates production of user-aligned tool-call trajectories under incomplete information.
  • The benchmark is built from privacy-sanitized real interaction traces.
  • It contains 10 user profiles, 36 tool sets, 1,065 turns, and 170 unique tools.
  • Task types include lack-of-information, single-tool, and multi-tool settings.
  • Experiments show current models still have difficulty with these tasks.

Entities

Sources