Trajectory Backdoor Attack: New Threat to Self-Evolving AI Skills
A recent paper published on arXiv (2608.08303) presents the Trajectory Backdoor Attack (TBA), which is a query-only method aimed at self-evolving skill systems within large language model (LLM) agents. These systems develop and refine skills based on execution trajectories, shifting skill acquisition from external sources to a secure internal process. Although this minimizes risks associated with direct skill injection, the authors contend that the evolution pipeline creates a new vulnerability: attackers can manipulate skill evolution by creating compromised trajectories through agent interactions. TBA exploits this weakness without needing to alter the skills directly. This research underscores increasing concerns regarding AI security as LLM agents gain autonomy and enhance self-evolving functions. The paper, authored by a team of researchers, is noted as a new submission on arXiv, likely submitted in 2026. It is pertinent to AI safety, cybersecurity, and the quest for reliable AI systems.
Key facts
- The paper is titled 'Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning'.
- It is available on arXiv with ID 2608.08303.
- The attack is called Trajectory Backdoor Attack (TBA).
- It is a query-only attack, meaning it does not require direct skill manipulation.
- Self-evolving skill systems automatically construct and update skills from execution trajectories.
- These systems shift skill acquisition from external marketplaces to a trusted evolution pipeline.
- The attack surface is introduced by the skill evolution pipeline.
- Attackers can indirectly steer skill evolution by inducing compromised trajectories through agent interactions.
Entities
Institutions
- arXiv