WuYuEval: New Benchmark Tests LLMs in Solid Waste Management
Researchers have introduced a new evaluation tool called WuYuEval, aimed at assessing large language models (LLMs) in solid waste management (SWM). Unlike existing assessments that focus mainly on broad knowledge, this benchmark zeroes in on decision-making in engineering, environmental, and policy contexts. WuYuEval includes two parts: a Foundation Module with 4,590 multiple-choice questions across six task types and eight domains, and an Expert Module with 247 scenario-based questions addressing multi-objective optimization and system design. The expert evaluation uses a unique scoring method with LLM-as-a-Judge and Elo comparisons. Testing 33 LLMs revealed notable performance differences, with the leading model showing outstanding results. You can find the research paper on arXiv under the ID 2608.07529.
Key facts
- WuYuEval is a multi-level benchmark for evaluating LLMs in solid waste management.
- It includes a Foundation Module with 4,590 closed-ended multiple-choice questions.
- The Foundation Module covers six task types and eight domain categories.
- An Expert Module contains 247 scenario-based open-ended questions.
- Expert tasks involve multi-objective optimization, constraint trade-offs, and system design.
- Evaluation uses anchor-calibrated LLM-as-a-Judge scoring and Elo-based pairwise comparison.
- 33 LLMs were tested, showing wide performance variation.
- The paper is available on arXiv (2608.07529).
Entities
Institutions
- arXiv