ARTFEED — Contemporary Art Intelligence

EcoAgent-Bench: New Benchmark for Budget-Conscious LLM Agents

ai-technology · 2026-08-07

A novel benchmark named EcoAgent-Bench has been launched to assess economic decision-making in large language model (LLM) agents while adhering to budget limitations. This benchmark is elaborated in a paper available on arXiv (2608.05519). Unlike conventional benchmarks that prioritize task completion, EcoAgent-Bench incorporates resource allocation as a core component of the task. Each task outlines priced actions and a defined budget, compelling agents to make strategic selections among local lookups, extensive searches, composite research tools, superior models, or human intervention. The benchmark features 304 tasks derived from real-world scenarios across five categories, adapted from GAIA, HotpotQA, and MuSiQue, focusing on four critical decisions: avoiding unnecessary escalation, escalating when local evidence is lacking, choosing a model tier, and halting on unsupported premises. The evaluation encompasses seven LLM agents in tool-API and workspace-CLI environments, along with four oracle scripted controls. Findings reveal that micro-averaged accuracy favors one-sided strategies, with always-escalate controls achieving high micro success yet struggling with save-oriented tasks. To tackle this issue, the authors suggest implementing an economic-consistency score, which reflects the lower accuracy on upgrades... (truncated for brevity)

Key facts

  • EcoAgent-Bench is a new benchmark for evaluating economic decision-making in LLM agents.
  • It introduces priced actions and explicit budgets for each task.
  • The benchmark includes 304 real-derived tasks from GAIA, HotpotQA, and MuSiQue.
  • It tests four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises.
  • Seven LLM agents are evaluated in tool-API and workspace-CLI settings.
  • Four oracle scripted controls are used for comparison.
  • Micro-averaged accuracy rewards one-sided policies like always-escalate.
  • An economic-consistency score is proposed as a better metric.
  • The paper is available on arXiv with ID 2608.05519.

Entities

Institutions

  • arXiv

Sources