Convergent Detour Hijacking: New Attack on LLM Agents' Skill Selection
A new research paper on arXiv (2608.12273v1) introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack targeting LLM agents that rely on third-party skills. The attack exploits the progressive-disclosure design of such agents, which use natural-language descriptions for skill selection and instruction bodies for planning. CDH couples two sequential control points: a static skill description establishes relevance during selection, while an aligned body fabricates plausible dependencies during planning. This allows an attacker-controlled coordinator to be attracted alongside legitimate skills, recruiting unnecessary benign skills into a bounded detour, and then re-en... (truncated due to source cut off). The paper highlights that prior work studied selection manipulation, malicious skill instructions, and tool-chain resource amplification separately, leaving their end-to-end composition unclear. CDH addresses this gap by combining these stages under shared semantic cover. The research is relevant to the security of AI systems, particularly as LLM agents become more prevalent in various applications.
Key facts
- Paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack on LLM agents.
- CDH targets LLM agents that use third-party skills with natural-language descriptions for selection and instruction bodies for planning.
- The attack exploits two sequential control points: skill description and instruction body.
- CDH couples selection manipulation and resource amplification, which were previously studied separately.
- The attack attracts an attacker-controlled coordinator alongside legitimate skills.
- It recruits unnecessary benign skills into a bounded detour.
- The paper is available on arXiv with ID 2608.12273v1.
- The research addresses a gap in understanding end-to-end composition of attacks on LLM agents.
Entities
Institutions
- arXiv