ARTFEED — Contemporary Art Intelligence

Systematic Review Reveals Gaps in LLM Agent Safety Guarantees

ai-technology · 2026-08-18

A recent systematic review encompassing 38 studies from 2022 to 2026 investigates the safety assurances for large language model (LLM) agents engaged in irreversible real-world tasks, including database modifications, API interactions, file handling, and autonomous tool usage. Adhering to PRISMA 2020 guidelines, this review sourced studies from six academic databases and highlights four major findings. A significant issue identified is the specification bottleneck, where natural-language-to-formal translation only achieves 24% to 35% semantic accuracy, hindering downstream verification. While runtime monitoring emerges as the most developed enforcement method, decreasing unsafe actions by 40% to 65% in controlled environments, the review emphasizes that no current system offers formally grounded, task-specific safety guarantees for the generated plans. Research remains disjointed across specification, verification, and enforcement, which restricts insights into the strengths and weaknesses of current methodologies. The review can be accessed on arXiv with the identifier 2608.14590.

Key facts

  • The review covers 38 studies published between 2022 and 2026.
  • Studies were retrieved from six academic databases.
  • The review follows the PRISMA 2020 systematic review guidelines.
  • Natural-language-to-formal translation achieves only 24% to 35% semantic correctness.
  • Runtime monitoring reduces unsafe actions by 40% to 65% in controlled settings.
  • No existing system provides formally grounded, task-level safety guarantees for LLM agent plans.
  • Research is fragmented across specification, verification, and enforcement.
  • The paper is available on arXiv with identifier 2608.14590.

Entities

Institutions

  • arXiv

Sources