ARTFEED — Contemporary Art Intelligence

Adversarial Persuasion: RL-Trained Agents Collapse LLM Accuracy to Near Zero

ai-technology · 2026-08-13

A recent study published on arXiv (2608.11624) indicates that large language models (LLMs) are significantly vulnerable to adversarial persuasion. The research shows that a single, targeted argument, even if incorrect, can drastically reduce their accuracy to nearly zero. This phenomenon is termed 'adversarial persuasion' and is examined through an adversarial reinforcement learning (RL) framework that trains persuader agents to alter a target model's response in just one interaction. The results reveal that RL-trained persuaders improve success rates from about 24% to over 93%, exposing weaknesses that traditional prompting overlooks. The authors emphasize that for LLMs to be reliable, they must resist harmful persuasion, highlighting an urgent need for stronger defenses against such tactics in AI systems.

Key facts

  • The study is published on arXiv with ID 2608.11624.
  • A single targeted persuasive argument can collapse model accuracy to near zero, even when the argument is factually false.
  • The threat is formalized as 'adversarial persuasion'.
  • An adversarial reinforcement learning framework trains persuader agents to change a target model's answer in a single interaction.
  • RL-trained persuaders raise persuasion success from approximately 24% to over 93%.
  • The research shows that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses.
  • LLMs increasingly debate, advise, and think collaboratively with humans and each other.
  • Resistance to harmful persuasion is identified as a core requirement for reliable behavior.

Entities

Institutions

  • arXiv

Sources