ARTFEED — Contemporary Art Intelligence

LLM-Generated Test Suites: Coverage and Mutation Scores May Not Predict Bug Detection

other · 2026-07-29

A new replicability study challenges the validity of using code coverage and mutation scores as proxy metrics for evaluating LLM-generated test suites. The research, conducted by a team of authors (not specified in the abstract), replicates prior work by Inozemtseva et al. and Papadakis et al., which found that for human-written tests, correlations between coverage, mutation, and real-bug detection largely vanish when test suite size is controlled. The study uses a wide range of test suites generated by diverse LLMs to re-examine these relationships. The findings raise concerns about the reliability of evaluations based on proxy metrics for LLM-generated tests, given that LLM-based test-generation workflows differ substantially from traditional approaches. The paper is available on arXiv under the ID 2607.22880.

Key facts

  • Study replicates Inozemtseva et al. and Papadakis et al. on LLM-generated tests
  • Examines correlations among coverage, mutation, and real-bug detection
  • Uses diverse LLMs and a wide range of test suites
  • Raises concerns about proxy metric validity for LLM-generated tests
  • Paper available on arXiv:2607.22880

Entities

Institutions

  • arXiv

Sources