ARTFEED — Contemporary Art Intelligence

E-Bench: Synthetic Benchmark for Multi-Step Tool-Use Agents

ai-technology · 2026-07-29

Researchers have introduced a new benchmark called E-Bench to evaluate large language models (LLMs) as agents that can perform multi-step tasks in changing settings. E-Bench consists of 323 tasks that create state changes across three different areas: Honor of Kings, QQ Music, and Tencent Meeting. By using graph-guided database filling, it effectively separates the creation of environments from task generation, allowing for reusable product environments without leftover components. The design features generator-solver asymmetry, presenting challenges that include both information and tool gaps, which require agents to uncover hidden data and make multiple tool calls before achieving a state change. This research aims to improve upon existing benchmarks that usually focus on isolated API calls or brief sequences. The full paper is accessible on arXiv under the reference 2607.23722.

Key facts

  • E-Bench is a fully synthetic benchmark for multi-step tool-use agents
  • Contains 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, Tencent Meeting
  • Uses graph-guided database filling for environment synthesis
  • Generator-solver asymmetry creates information and tool gaps
  • Outcomes are graded deterministically
  • Addresses limitations of existing benchmarks focusing on isolated API calls or short trajectories
  • Paper available on arXiv: 2607.23722

Entities

Institutions

  • arXiv
  • Honor of Kings
  • QQ Music
  • Tencent Meeting

Sources