InteractComp: Benchmark for Search Agents with Ambiguous Queries
A new benchmark called InteractComp has been developed by researchers to assess the capability of search agents in identifying and clarifying ambiguous queries through interaction. This benchmark tackles a common issue where user inquiries are vague or lacking detail, necessitating clarification from the agents. It features 210 meticulously curated questions spanning 9 different domains, employing a target-distractor approach to create controlled ambiguity that can only be resolved through interaction. The evaluation involved 17 models, highlighting notable deficiencies in the current agents' performance in these situations. This research is available on arXiv with the identifier 2510.24668.
Key facts
- InteractComp is a benchmark for evaluating search agents on ambiguous queries.
- It tests whether agents can recognize ambiguity and interact to resolve it.
- The benchmark includes 210 expert-curated questions across 9 domains.
- Questions are created using a target-distractor methodology.
- Ambiguity is controlled and resolvable only through interaction.
- 17 models were evaluated in the study.
- Results show current agents struggle with ambiguous queries.
- The paper is available on arXiv with ID 2510.24668.
Entities
Institutions
- arXiv