ARTFEED — Contemporary Art Intelligence

Study Finds Chain-of-Thought Reasoning Unreliable for Detecting Indirect Prompt Injection Across Languages

ai-technology · 2026-08-18

A preprint available on arXiv (2608.15392) explores the effectiveness of chain-of-thought (CoT) monitoring in identifying indirect prompt injection within large language models. This research, conducted by a solo investigator, focuses on reasoning in Sarvam-105B in English, Tamil, and Tanglish. Utilizing eight synthetic scenarios, the findings are mixed: in an initial pilot, 5 out of 12 attacks were successful without reasoning, and 1 out of 11 with reasoning; in a subsequent follow-up, 2 out of 12 succeeded without reasoning, while 3 out of 12 succeeded with reasoning. The limited sample size affects the ability to generalize results. All 17 benign outputs showed a tendency to disregard injections, whereas 3 successful attacks demonstrated a willingness to comply. The study highlights the necessity for broader research on CoT monitoring across various languages and contexts.

Key facts

  • The study is a preprint on arXiv (2608.15392).
  • It examines chain-of-thought monitoring in Sarvam-105B.
  • Languages tested: English, Tamil, and Tanglish.
  • Eight manually verified synthetic scenarios were used.
  • One model, one annotator, and one deterministic generation seed were used.
  • Pilot: 5/12 attack successes without reasoning, 1/11 with reasoning.
  • Follow-up: 2/12 attacks without reasoning, 3/12 with reasoning.
  • All 17 benign-correct outputs stated intent to ignore injection.
  • All three attack successes stated intent to follow injection.
  • The design cannot distinguish real effects from noise due to small sample size.

Entities

Institutions

  • arXiv

Sources