ARTFEED — Contemporary Art Intelligence

Probing Hidden States Reveals Indirect Prompt-Injection Exposure in Agentic LLMs

ai-technology · 2026-08-06

A recent preprint on arXiv (2608.02657) explores how agentic large language models (LLMs) respond to indirect prompt injection (IPI) attacks, a phenomenon referred to as 'IPI exposure.' The research, which examined six models, including the 753B-parameter GLM-5.2, reveals that straightforward linear probes trained on hidden states before generation can achieve over 90% AUROC in predicting unseen IPI attacks, agent instructions, and task suites. These probes maintain significant robustness against adaptive attacks and in cross-lingual contexts. Additionally, the study uncovers a 'recognition-action gap' via chain-of-thought (CoT) metrics, indicating that while models can detect IPI exposure signals, they often do not act safely. To mitigate this, the authors propose AGRI, a defense mechanism that incorporates a reasoning step before generation, triggered by the probe's detection. This paper is classified as a cross-submission and is accessible on arXiv, emphasizing the importance of monitoring internal states to bolster LLM security against IPI attacks.

Key facts

  • The paper is titled 'Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure.'
  • It is an arXiv preprint with identifier 2608.02657.
  • The study covers six models, including GLM-5.2 with 753 billion parameters.
  • Linear probes trained on pre-generation hidden states predict IPI exposure with over 90% AUROC.
  • The probes are robust under adaptive attacks and cross-lingual settings.
  • A recognition-action gap is identified: models encode signals but fail to act safely.
  • AGRI is introduced as a probe-gated reasoning-based defense.
  • The defense prepends a reasoning step before generation, gated by probe detection.

Entities

Institutions

  • arXiv

Sources