ARTFEED — Contemporary Art Intelligence

AcquaBench Audits Success Provenance in AI Agent Evaluation

ai-technology · 2026-07-29

A recent preprint on arXiv (2607.24054) presents AcquaBench, a benchmarking tool aimed at evaluating the success provenance of agents. The researchers contend that correct responses may mask whether an agent's success stemmed from intended reasoning or from obtaining relevant information during the evaluation process. AcquaBench employs matched value substitutions—CLEAN, GOLD, and SHAM—across four standardized surfaces, utilizing joint qid-clustered analysis. CLEAN maintains benchmark-approved data, GOLD provides the correct target, and SHAM keeps the source structure while replacing it with an incorrect value. The score difference between GOLD and CLEAN assesses the impact of correct-target availability, while the difference between GOLD and SHAM examines if this response reflects target correctness beyond mere exposure. In experiment D0, GOLD outperformed SHAM by 19.1 to 25.9 percentage points, suggesting that success is often contingent on target availability rather than reasoning.

Key facts

  • arXiv preprint 2607.24054 introduces AcquaBench
  • AcquaBench audits success provenance in agent evaluation
  • Uses CLEAN, GOLD, and SHAM value substitutions
  • Four standardized surfaces with joint qid-clustered analysis
  • GOLD minus CLEAN measures score response to correct-target availability
  • GOLD minus SHAM tests if response tracks target correctness
  • In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points
  • Success often depends on target availability rather than reasoning

Entities

Institutions

  • arXiv

Sources