WebCoderBench: A New Benchmark for LLM-Generated Web Apps
A new benchmark called WebCoderBench has been launched by researchers to assess large language models (LLMs) in their ability to generate web applications. This benchmark is derived from 1,572 actual user requirements, encompassing various modalities and expression styles to accurately mirror user intentions. It features 24 detailed evaluation metrics across 9 different perspectives, integrating both rule-based methods and LLM-as-a-judge approaches for a fully automated, objective evaluation. WebCoderBench tackles the difficulties of assessing LLM-generated web applications, including the necessity for real-world user requirements and generalizable metrics without ground-truth implementations. This pioneering benchmark is documented in a paper on arXiv (arXiv:2601.02430) under the announcement type 'replace-cross', aiming to enhance code generation evaluation.
Key facts
- WebCoderBench is a new benchmark for web app generation by LLMs.
- It comprises 1,572 real user requirements.
- It provides 24 fine-grained evaluation metrics across 9 perspectives.
- Evaluation combines rule-based and LLM-as-a-judge paradigms.
- It is fully automated, objective, and general.
- It is the first real-world-collected, generalizable, and interpretable benchmark for web app generation.
- The paper is available on arXiv with ID 2601.02430.
- The announcement type is 'replace-cross'.
Entities
Institutions
- arXiv