ARTFEED — Contemporary Art Intelligence

LLM Code Generation Vulnerabilities Exceed 56% in Realistic Scenarios

ai-technology · 2026-07-29

A new study from arXiv (2607.23088) reveals that large language models (LLMs) used for code generation exhibit average vulnerability rates exceeding 56% across realistic risk scenarios. The research identifies three key scenarios leading to security flaws: ambiguous requirements, under-specified operational context, and security-functionality conflict. The benchmark includes 2,700 test cases for fine-grained evaluation. Eight state-of-the-art LLMs were tested, all showing high vulnerability rates.

Key facts

  • arXiv paper 2607.23088
  • LLMs for code generation have over 56% average vulnerability rates
  • Three risk scenarios: ambiguous requirements, under-specified operational context, security-functionality conflict
  • Benchmark of 2,700 test cases
  • Eight state-of-the-art LLMs evaluated
  • Existing benchmarks fail to capture real-world ambiguous prompts

Entities

Institutions

  • arXiv

Sources