WeSCE: Benchmarking Security Drift in LLM-Driven Code Editing
A new benchmark named WeSCE has been developed by researchers to assess security drift in code modifications made by large language models (LLMs) operating under weak-security conditions, where only functional goals are outlined without specific security criteria. This benchmark includes 400 executable programs sourced from actual code, focusing on tasks like feature addition, feature removal, bug fixing, and refactoring. To evaluate security drift, the researchers suggest a continuous risk representation that integrates various vulnerability signals into a single framework. They establish drift metrics that reflect shifts in overall risk, extreme severity, and vulnerability distribution amid code alterations, offering insights into security from average to worst-case scenarios. This initiative aims to tackle the rising issue of potential security vulnerabilities introduced by LLM-assisted code editing. Developers and researchers can utilize this benchmark to evaluate and reduce such risks. The paper can be found on arXiv in the Computer Science > Cryptography and Security section.
Key facts
- WeSCE is a benchmark for measuring security drift in LLM-driven code editing.
- It consists of 400 executable programs derived from real-world code.
- Tasks include feature addition, feature removal, bug fixing, and refactoring.
- A continuous risk representation aggregates heterogeneous vulnerability signals.
- Drift measures capture changes in overall risk, worst-case severity, and vulnerability distribution.
- The benchmark provides a multi-scale view of security from average-case to worst-case.
- The paper is categorized under Computer Science > Cryptography and Security.
- The arXiv ID is 2608.15092.
Entities
Institutions
- arXiv