ARTFEED — Contemporary Art Intelligence

LLM System Instruction Leakage via Encoding Attacks: Evaluation Framework

ai-technology · 2026-08-13

An arXiv paper (2604.01039v4) presents a framework for automated evaluation designed to assess the confidentiality of system instructions in Large Language Models (LLMs) when extraction requests are reformulated as structured output or encoding tasks. System instructions play a crucial role in implementing safety protocols, shaping agent behavior, and safeguarding sensitive operational contexts in AI applications, often containing critical data like API keys, internal guidelines, and confidential workflow definitions. The paper identifies the leakage of system instructions as a significant security threat, as highlighted in the OWASP Top 10 for LLM Applications. The framework tests seven prevalent models and 46 validated system instructions to evaluate their susceptibility to encoding-based extraction attacks, aiming to enhance LLM security strategies.

Key facts

  • arXiv paper 2604.01039v4 introduces an automated evaluation framework for testing LLM system instruction confidentiality.
  • The framework tests extraction via encoding or structured output tasks.
  • System instructions may contain sensitive data like API credentials, internal policies, and privileged workflow definitions.
  • System instruction leakage is a critical security risk highlighted in OWASP Top 10 for LLM Applications.
  • Many LLM applications rely on refusal-based instructions to block direct requests.
  • The framework evaluates seven common models and 46 verified system instructions.
  • The paper aims to inform hardening strategies for LLM system instructions.
  • The paper is available on arXiv with URL https://arxiv.org/abs/2604.01039.

Entities

Institutions

  • arXiv
  • OWASP

Sources