LLM Code Generation Vulnerabilities Exceed 56% in Realistic Scenarios
A new study from arXiv (2607.23088) reveals that large language models (LLMs) used for code generation exhibit average vulnerability rates exceeding 56% across realistic risk scenarios. The research identifies three key scenarios leading to security flaws: ambiguous requirements, under-specified operational context, and security-functionality conflict. The benchmark includes 2,700 test cases for fine-grained evaluation. Eight state-of-the-art LLMs were tested, all showing high vulnerability rates.
Key facts
- arXiv paper 2607.23088
- LLMs for code generation have over 56% average vulnerability rates
- Three risk scenarios: ambiguous requirements, under-specified operational context, security-functionality conflict
- Benchmark of 2,700 test cases
- Eight state-of-the-art LLMs evaluated
- Existing benchmarks fail to capture real-world ambiguous prompts
Entities
Institutions
- arXiv