Probing Hidden States Reveals Indirect Prompt-Injection Exposure in Agentic LLMs
A recent preprint on arXiv (2608.02657) explores how agentic large language models (LLMs) respond to indirect prompt injection (IPI) attacks, a phenomenon referred to as 'IPI exposure.' The research, which examined six models, including the 753B-parameter GLM-5.2, reveals that straightforward linear probes trained on hidden states before generation can achieve over 90% AUROC in predicting unseen IPI attacks, agent instructions, and task suites. These probes maintain significant robustness against adaptive attacks and in cross-lingual contexts. Additionally, the study uncovers a 'recognition-action gap' via chain-of-thought (CoT) metrics, indicating that while models can detect IPI exposure signals, they often do not act safely. To mitigate this, the authors propose AGRI, a defense mechanism that incorporates a reasoning step before generation, triggered by the probe's detection. This paper is classified as a cross-submission and is accessible on arXiv, emphasizing the importance of monitoring internal states to bolster LLM security against IPI attacks.
Key facts
- The paper is titled 'Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure.'
- It is an arXiv preprint with identifier 2608.02657.
- The study covers six models, including GLM-5.2 with 753 billion parameters.
- Linear probes trained on pre-generation hidden states predict IPI exposure with over 90% AUROC.
- The probes are robust under adaptive attacks and cross-lingual settings.
- A recognition-action gap is identified: models encode signals but fail to act safely.
- AGRI is introduced as a probe-gated reasoning-based defense.
- The defense prepends a reasoning step before generation, gated by probe detection.
Entities
Institutions
- arXiv