ARTFEED — Contemporary Art Intelligence

LLMs Show Promise but Fall Short in Security Protocol Analysis

ai-technology · 2026-07-29

A recent study investigated the capability of Large Language Models (LLMs) to conduct symbolic security protocol analysis, comparing their performance to established formal tools like ProVerif and OFMC. Researchers examined GPT and DeepSeek in both chat and reasoning modes across three trials involving 130 obfuscated AnB/AnBx protocols, addressing 388 security objectives. The chat models demonstrated a recall rate between 69-81%, yet their precision fell below 31%. Conversely, reasoning models improved precision, achieving 66.5% for GPT and 45.4% for DeepSeek, although they identified just over half of the attacks. Notably, DeepSeek's two modes utilize a single underlying model, enhancing precision from 27.2% to 45.4%. All models struggled most with authentication goals, with reasoning models detecting fewer than half of the attacks.

Key facts

  • Study evaluates LLMs for symbolic security protocol analysis
  • Tests GPT and DeepSeek in chat and reasoning modes
  • Three runs on 130 obfuscated AnB/AnBx protocols
  • Covers 388 security goals
  • Scored against ProVerif and OFMC
  • Chat models: 69-81% recall, precision below 31%
  • Reasoning models: GPT 66.5% precision, DeepSeek 45.4% precision
  • All models worst on authentication goals

Entities

Institutions

  • arXiv

Sources