LLM-Generated Test Suites: Coverage and Mutation Scores May Not Predict Bug Detection
A new replicability study challenges the validity of using code coverage and mutation scores as proxy metrics for evaluating LLM-generated test suites. The research, conducted by a team of authors (not specified in the abstract), replicates prior work by Inozemtseva et al. and Papadakis et al., which found that for human-written tests, correlations between coverage, mutation, and real-bug detection largely vanish when test suite size is controlled. The study uses a wide range of test suites generated by diverse LLMs to re-examine these relationships. The findings raise concerns about the reliability of evaluations based on proxy metrics for LLM-generated tests, given that LLM-based test-generation workflows differ substantially from traditional approaches. The paper is available on arXiv under the ID 2607.22880.
Key facts
- Study replicates Inozemtseva et al. and Papadakis et al. on LLM-generated tests
- Examines correlations among coverage, mutation, and real-bug detection
- Uses diverse LLMs and a wide range of test suites
- Raises concerns about proxy metric validity for LLM-generated tests
- Paper available on arXiv:2607.22880
Entities
Institutions
- arXiv