Search papers, labs, and topics across Lattice.
This paper introduces VAKRA, a comprehensive benchmark designed to evaluate multi-hop reasoning capabilities of agents interacting with structured APIs and document collections across diverse domains. The study reveals that even state-of-the-art models struggle significantly, achieving only 70.4% accuracy on simpler tasks and plummeting to as low as 2.4% on complex, policy-constrained queries. Trace analysis indicates that the primary challenges lie in language-mediated reasoning tasks, such as entity disambiguation and cross-source grounding, rather than in the mechanics of tool invocation.
Multi-hop reasoning in API interactions is a critical bottleneck, with top models faltering under policy constraints and complex queries.
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA