ARTFEED — Contemporary Art Intelligence

Study Diagnoses Search Behavior and Failure Modes in Long-Horizon Search Agents

ai-technology · 2026-08-04

A recent arXiv preprint (2608.01913) introduces a trajectory-level assessment of long-horizon search agents that tackle challenging information-seeking inquiries through iterative search queries. The research evaluates evidence gathered at each stage using human-annotated document-level relevance assessments, distinguishing between two phases: the evidence retrieved by an agent and its effectiveness in utilizing that evidence. This separation enables the identification of failures as either retrieval gaps (missing necessary evidence) or utilization gaps (relevant evidence not used properly). Keeping the retrieval model and evaluation framework constant, the authors analyze six agents on BrowseComp-Plus and confirm their results on BrowseComp via an open-web search API. The study seeks to understand how increased search effort correlates with improved answers, shedding light on agent behavior and failure mechanisms.

Key facts

  • The study is a trajectory-level diagnosis of long-harizon search agents.
  • It uses human-annotated document-level relevance judgments.
  • Failures are decomposed into retrieval gaps and utilization gaps.
  • Six agents are compared on BrowseComp-Plus.
  • Validation is done on BrowseComp with an open-web search API.
  • The retrieval model and evaluation harness are held fixed.
  • The paper is available on arXiv with ID 2608.01913.

Entities

Institutions

  • arXiv

Sources