Study Finds Chain-of-Thought Reasoning Unreliable for Detecting Indirect Prompt Injection Across Languages
A preprint available on arXiv (2608.15392) explores the effectiveness of chain-of-thought (CoT) monitoring in identifying indirect prompt injection within large language models. This research, conducted by a solo investigator, focuses on reasoning in Sarvam-105B in English, Tamil, and Tanglish. Utilizing eight synthetic scenarios, the findings are mixed: in an initial pilot, 5 out of 12 attacks were successful without reasoning, and 1 out of 11 with reasoning; in a subsequent follow-up, 2 out of 12 succeeded without reasoning, while 3 out of 12 succeeded with reasoning. The limited sample size affects the ability to generalize results. All 17 benign outputs showed a tendency to disregard injections, whereas 3 successful attacks demonstrated a willingness to comply. The study highlights the necessity for broader research on CoT monitoring across various languages and contexts.
Key facts
- The study is a preprint on arXiv (2608.15392).
- It examines chain-of-thought monitoring in Sarvam-105B.
- Languages tested: English, Tamil, and Tanglish.
- Eight manually verified synthetic scenarios were used.
- One model, one annotator, and one deterministic generation seed were used.
- Pilot: 5/12 attack successes without reasoning, 1/11 with reasoning.
- Follow-up: 2/12 attacks without reasoning, 3/12 with reasoning.
- All 17 benign-correct outputs stated intent to ignore injection.
- All three attack successes stated intent to follow injection.
- The design cannot distinguish real effects from noise due to small sample size.
Entities
Institutions
- arXiv