ARTFEED — Contemporary Art Intelligence

Validity Audit of Agent-Safety Benchmarks Reveals Scoring Flaws

ai-technology · 2026-08-03

An arXiv paper (2607.28685) has conducted a validity audit on four benchmarks related to agent safety: R-Judge, InjecAgent, AgentHarm, and AgentDojo. The authors contend that the scores from these benchmarks are often used interchangeably to assess an agent's safety, even though they evaluate distinct behaviors. They validate the benchmarks by applying each one under its official implementation and using the scorer provided by the authors across 22 models. Additionally, they assess MMLU and GPQA under a unified protocol as a composite of capabilities. The study highlights a significant issue with metrics: for any binary trace-judgment benchmark evaluated by F1, an 'always positive' policy achieves F1 = 2π/(1+π). For R-Judge, this results in 0.690, surpassing five of the 21 models that effectively discriminate. The three broad-coverage benchmarks rank the same 18 models differently, revealing a small-panel artifact in their trade-off: R-Judge specificity correlates with AgentHarm safety at -0.64 for n=7 and +0.02 for n=18, with a quarter of random size-7 subsets reaching |ρ| ≥ 0.5 around that near-zero correlation. The authors conclude that existing agent-safety benchmarks may not provide a reliable measure of safety, urging caution in interpreting their scores.

Key facts

  • The paper is titled 'Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks'.
  • It is available on arXiv with ID 2607.28685.
  • Four benchmarks are audited: R-Judge, InjecAgent, AgentHarm, and AgentDojo.
  • Up to 22 models are tested under official implementations and author-provided scorers.
  • MMLU and GPQA are measured under one protocol as a capability composite.
  • An 'always positive' policy achieves F1 = 2π/(1+π) on binary trace-judgment benchmarks.
  • On R-Judge, that F1 is 0.690, above five of 21 models that discriminate.
  • The three broad-coverage benchmarks rank 18 models differently.
  • R-Judge specificity vs AgentHarm safety correlation is -0.64 at n=7 and +0.02 at n=18.
  • A quarter of random size-7 subsets reach |ρ| ≥ 0.5 around that near-zero correlation.

Entities

Institutions

  • arXiv

Sources