EcoAgent-Bench: New Benchmark for Budget-Conscious LLM Agents
A novel benchmark named EcoAgent-Bench has been launched to assess economic decision-making in large language model (LLM) agents while adhering to budget limitations. This benchmark is elaborated in a paper available on arXiv (2608.05519). Unlike conventional benchmarks that prioritize task completion, EcoAgent-Bench incorporates resource allocation as a core component of the task. Each task outlines priced actions and a defined budget, compelling agents to make strategic selections among local lookups, extensive searches, composite research tools, superior models, or human intervention. The benchmark features 304 tasks derived from real-world scenarios across five categories, adapted from GAIA, HotpotQA, and MuSiQue, focusing on four critical decisions: avoiding unnecessary escalation, escalating when local evidence is lacking, choosing a model tier, and halting on unsupported premises. The evaluation encompasses seven LLM agents in tool-API and workspace-CLI environments, along with four oracle scripted controls. Findings reveal that micro-averaged accuracy favors one-sided strategies, with always-escalate controls achieving high micro success yet struggling with save-oriented tasks. To tackle this issue, the authors suggest implementing an economic-consistency score, which reflects the lower accuracy on upgrades... (truncated for brevity)
Key facts
- EcoAgent-Bench is a new benchmark for evaluating economic decision-making in LLM agents.
- It introduces priced actions and explicit budgets for each task.
- The benchmark includes 304 real-derived tasks from GAIA, HotpotQA, and MuSiQue.
- It tests four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises.
- Seven LLM agents are evaluated in tool-API and workspace-CLI settings.
- Four oracle scripted controls are used for comparison.
- Micro-averaged accuracy rewards one-sided policies like always-escalate.
- An economic-consistency score is proposed as a better metric.
- The paper is available on arXiv with ID 2608.05519.
Entities
Institutions
- arXiv