ARTFEED — Contemporary Art Intelligence

HarnessRisk: New Lifecycle Benchmark for AI Agent Safety

ai-technology · 2026-08-19

A preprint titled "HarnessRisk" has been published on arXiv (2608.17597), introducing a lifecycle-focused benchmark for the safety of agent harnesses. This benchmark evaluates large language model agent harnesses through six distinct operational phases: configuration of the harness, extension of capabilities, operation during runtime, persistence of state, control of actions, and recovery from incidents. It includes 128 sandboxed scenarios, each combining a benign user goal with an adversarial directive within an untrustworthy workflow artifact. Evaluations are based on utility, success rate of attacks, persistence, and detection. The framework tests three harnesses, six language models, and 14 configurations, identifying safety failures in tools, extensions, state management, permissions, and external actions, thereby filling gaps in current safety benchmarks that often focus on single attack types or limited operational contexts.

Key facts

  • HarnessRisk is a new lifecycle-oriented benchmark for agent harness safety.
  • It was released on arXiv as preprint 2608.17597.
  • The benchmark covers six operational phases.
  • It includes 128 sandboxed test cases.
  • Each case pairs a benign objective with an adversarial instruction.
  • Four evaluation metrics are used: utility, attack success rate, persistence, and detection.
  • Evaluations covered three harnesses, six language models, and 14 configurations.
  • Existing safety benchmarks target individual attack mechanisms or limited operational settings.

Entities

Sources