ARTFEED — Contemporary Art Intelligence

PoisonedEvolution: Trajectory Poisoning Attack on Self-Evolving Agent Skill Systems

ai-technology · 2026-08-07

A recent study presents PoisonedEvolution, an attack method that targets self-evolving skill (SES) systems by manipulating agent trajectories into long-lasting skills. This method takes advantage of how skills transition from temporary to enduring instructions. The attacker operates as a skill-visible black-box entity, capable of examining a specific skill and providing limited evidence, but lacking access to private resources or the ability to modify the skill bank. The attack hinges on three components: Inclusion, Evolution Attribution, and Realization, with Attribution being the critical limiting factor. Researchers tested four security-effect families using inert canary specifications. With 10% support from the attacker, PoisonedEvolution successfully integrated target behaviors in 546 out of 600 trials across six LLM evolvers, resulting in a 91.0% success rate. The paper can be found on arXiv (identifier 2608.05563).

Key facts

  • PoisonedEvolution is a trajectory-poisoning attack on self-evolving skill (SES) systems.
  • SES systems distill agent trajectories into persistent skills, turning untrusted experience into trusted instruction.
  • The attacker is skill-visible black-box: can inspect target skill and contribute bounded evidence, but cannot observe private pools or evolution logic, or edit the skill bank.
  • Attack requires Inclusion, Evolution Attribution, and Realization; Attribution is the distinctive bottleneck.
  • At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER).
  • On the Trace2Skill pipeline at the same ratio, it also embeds target behaviors.
  • The paper is on arXiv with ID 2608.05563.
  • The paper is a cross-type announcement.

Entities

Institutions

  • arXiv

Sources