PoisonedEvolution: Trajectory Poisoning Attack on Self-Evolving Agent Skill Systems
A recent study presents PoisonedEvolution, an attack method that targets self-evolving skill (SES) systems by manipulating agent trajectories into long-lasting skills. This method takes advantage of how skills transition from temporary to enduring instructions. The attacker operates as a skill-visible black-box entity, capable of examining a specific skill and providing limited evidence, but lacking access to private resources or the ability to modify the skill bank. The attack hinges on three components: Inclusion, Evolution Attribution, and Realization, with Attribution being the critical limiting factor. Researchers tested four security-effect families using inert canary specifications. With 10% support from the attacker, PoisonedEvolution successfully integrated target behaviors in 546 out of 600 trials across six LLM evolvers, resulting in a 91.0% success rate. The paper can be found on arXiv (identifier 2608.05563).
Key facts
- PoisonedEvolution is a trajectory-poisoning attack on self-evolving skill (SES) systems.
- SES systems distill agent trajectories into persistent skills, turning untrusted experience into trusted instruction.
- The attacker is skill-visible black-box: can inspect target skill and contribute bounded evidence, but cannot observe private pools or evolution logic, or edit the skill bank.
- Attack requires Inclusion, Evolution Attribution, and Realization; Attribution is the distinctive bottleneck.
- At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER).
- On the Trace2Skill pipeline at the same ratio, it also embeds target behaviors.
- The paper is on arXiv with ID 2608.05563.
- The paper is a cross-type announcement.
Entities
Institutions
- arXiv