ARTFEED — Contemporary Art Intelligence

Study Evaluates LLM Defenders Against AI-Generated Social Engineering in Live Interactions

ai-technology · 2026-08-13

A recent paper published on arXiv (2608.10239) explores how effective large language model (LLM)-based defense mechanisms are in safeguarding users from AI-driven social engineering attacks during real-time, turn-by-turn exchanges. The research introduces 'trust-chain localization,' a framework that assesses whether an interaction fails due to actor authority, asset control, verification adequacy, or transaction pathway issues. The team created a controlled dataset comprising 300 online housing scenarios across 20 families, including legitimate examples, four types of structural failures, and three surface conditions. Five defender models were tested on this dataset in both stateful and one-shot static formats, leading to 1,500 evaluations per protocol, totaling 3,000. While no model exhibited explicit unsafe compliance, defensive effectiveness showed significant variation, with intervention rates between 0% and 96.3%. The findings emphasize the importance of identifying structural risk sources rather than merely responding to surface indicators.

Key facts

  • Paper on arXiv: 2608.10239
  • Focuses on LLM-based defenders against AI-generated social engineering
  • Introduces 'trust-chain localization' concept
  • Corpus of 300 online-housing cases across 20 scenario families
  • Five defender models evaluated
  • 1,500 model-case evaluations per protocol, 3,000 total
  • Intervention rates ranged from 0% to 96.3%
  • No model produced explicit unsafe compliance

Entities

Institutions

  • arXiv

Sources