VAKRA: New Benchmark Evaluates Multi-Hop Reasoning Across APIs and Retrieval
A new benchmark called VAKRA (Valuating API and Knowledge Retrieval Agents) has been developed to assess the abilities of AI agents in business environments where reasoning across structured APIs and document collections is essential. Unlike existing benchmarks that evaluate these skills separately, VAKRA integrates them, offering over 8,000 executable APIs from 62 different domains. It presents tasks in three progressively challenging scenarios: varied API interaction methods, multi-hop reasoning with structured APIs, and reasoning from multiple sources under natural-language tool-use policy constraints. The accuracy of predictions is confirmed by re-executing tool calls on live APIs, allowing for various valid outcomes. Evaluations of frontier and open-weight models using a fixed ReAct harness show that the top model only reaches 70.4% accuracy on single-hop endpoint tasks and falls to 50–51% for compositional APIs, highlighting the benchmark's goal of realistically measuring agent performance in complex, real-world situations.
Key facts
- VAKRA is a benchmark for evaluating AI agents in enterprise settings.
- It includes over 8,000 executable APIs across 62 domains.
- Tasks span three settings: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints.
- Correctness is verified by re-executing predicted tool calls against live APIs.
- A fixed ReAct harness is used to isolate model capabilities from agent architecture.
- The best model achieves 70.4% on single-hop endpoint-style tasks.
- Performance drops to 50–51% on compositional APIs.
- The benchmark is introduced in a paper on arXiv (arXiv:2608.12282).
Entities
Institutions
- arXiv
- IBM
- IBM Research
- GitHub
- Hugging Face
- OpenAI
- Google Analytics
- Tableau
- UC Berkeley