MerchantBench: New Benchmark for LLM Agents in E-Commerce Operations
MerchantBench, a novel benchmark, has been launched to assess the long-term consistency of large language model (LLM) agents in e-commerce settings. This benchmark is based on 98,843 authentic e-commerce product records and replicates a 365-day order-level scenario. It offers 26 tools for agent engagement, encompassing tasks like product sourcing, pricing control, cash-flow management, and adapting to mixed-latency feedback. The primary aim is to evaluate agents' capabilities to sustain purposeful actions over extended periods while responding to accumulated data, which is vital for practical applications. MerchantBench also addresses the shortcomings of current benchmarks that prioritize limited tasks with immediate success metrics. Detailed information can be found in a paper on arXiv (arXiv:2607.28956).
Key facts
- MerchantBench is a new benchmark for LLM agents in e-commerce.
- It is grounded in 98,843 real e-commerce product records.
- The simulation runs for 365 days at order-level granularity.
- It provides 26 tools for agent interaction.
- Tasks include product sourcing, listing and pricing control, cash-flow management, and mixed-latency feedback adaptation.
- The benchmark focuses on long-term coherence, not just immediate success.
- The paper is available on arXiv with identifier 2607.28956.
Entities
Institutions
- arXiv