ARTFEED — Contemporary Art Intelligence

SHE: A New Framework for Evolving LLM Agent Safety Harnesses

ai-technology · 2026-08-11

A new paper called 'SHE: Trajectory-driven Safety Harness Evolution for LLM Agents' has been published on arXiv with the ID 2608.09885. It examines how to improve safety for large language model agents, arguing that safety isn't just about the model's weights—it's also about the harness that manages context, memory, tools, permissions, and control during runtime. The authors note that existing safety methods treat the harness as a fixed part, which limits its ability to adapt to fresh risks. They propose a framework named Safety Harness Evolution (SHE) that identifies safe boundaries from rollout trajectories and divides the harness into four key components: System Prompt, Rule Bank, Safety Memory, and Tool Policy, each with its own safety function. The paper also outlines an evolution loop to help diagnose and address failures, allowing for specific updates and continuous improvement of the harness. This research is classified as a new announcement and is available through the provided arXiv link.

Key facts

  • Paper title: 'SHE: Trajectory-driven Safety Harness Evolution for LLM Agents'
  • Published on arXiv with ID 2608.09885
  • Proposes Safety Harness Evolution (SHE) framework
  • Decomposes harness into System Prompt, Rule Bank, Safety Memory, and Tool Policy
  • Uses attribution-guided evolution loop from rollout trajectories
  • Addresses limitations of fixed deployment artifacts in safety mechanisms
  • Focuses on LLM agent safety beyond model weights
  • Available at https://arxiv.org/abs/2608.09885

Entities

Institutions

  • arXiv

Sources