ARTFEED — Contemporary Art Intelligence

SearchAuditor: New Framework to Audit and Repair Failures in Long-Horizon Search Agents

ai-technology · 2026-08-07

A new study has introduced SearchAuditBench, a tool created to evaluate how well large language model auditors can spot, attribute, and fix errors in long-horizon search agents. This benchmark consists of 1,243 failed attempts, averaging 73.1 messages and 65.1K tokens, drawn from eight open-weight models across five deep-search benchmarks. Experts have thoroughly annotated each attempt, marking the key error step, its search-related cause, and a possible fix, along with grading criteria. The study also presents SearchAuditor, a comprehensive framework that efficiently identifies these failures, eliminating the need for manual checks of lengthy execution logs. It reveals that even small reasoning errors in deep search agents can result in seemingly fluent but incorrect answers. The research can be found on arXiv under the ID 2608.05212.

Key facts

  • SearchAuditBench includes 1,243 failed trajectories.
  • Trajectories average 73.1 messages and 65.1K tokens.
  • Data collected from eight open-weight models on five deep-search benchmarks.
  • Each trajectory is expert-annotated with critical error step, root cause, and reference repair.
  • SearchAuditor is a multi-perspective auditing framework.
  • The paper is available on arXiv (2608.05212).
  • The research aims to reduce human burden in diagnosing search agent failures.
  • Long-horizon search agents are complex and fragile.

Entities

Institutions

  • arXiv

Sources