ComboShoppingBench: New Benchmark for Budget-Constrained Basket Shopping with Coupons
ComboShoppingBench, a novel benchmark for agentic shopping, has been introduced in a paper available on arXiv (2608.09282). This benchmark is aimed at assessing LLM agents engaged in budget-limited basket shopping that incorporates coupons. It tackles the difficulty of assembling a collection of complementary items rather than just a single product, a common requirement in scenarios like meal preparation, event planning, device setup, and group takeout orders. These situations necessitate comprehensive reasoning regarding item compatibility, availability, delivery fees, store requirements, and budget constraints. The authors emphasize the evaluation's complexity, as multiple baskets can fulfill the same request, rendering exact-match metrics ineffective. ComboShoppingBench offers a verifiable environment for basket creation within a simulated commerce context, enabling agents to tackle intricate multi-item shopping tasks.
Key facts
- ComboShoppingBench is introduced in a paper on arXiv with ID 2608.09282.
- The benchmark evaluates LLM agents for budget-constrained basket shopping with coupons.
- It focuses on constructing baskets of complementary items, not single products.
- Real-world scenarios include device setup, meal preparation, event planning, and group takeout ordering.
- Tasks require reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets.
- Exact-match metrics are unsuitable because multiple baskets may satisfy the same request.
- Semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments.
- The benchmark uses a simulated commerce and takeout environment.
- An exploration agent constructs a feasible and semantically coherent basket as a witness.
- The witness guides the generation of coupons and budget constraints.
Entities
Institutions
- arXiv