Search papers, labs, and topics across Lattice.
This study introduces DeepStress, a novel stress-testing framework designed to evaluate the robustness of deep search agents against unreliable evidence by simulating a controlled synthetic environment. By systematically manipulating the trustworthiness, relevance, and factuality of the evidence presented, the authors reveal significant disparities in how various search agents perform under challenging conditions. The findings highlight the need for improved metrics to assess not only the outcomes of document systems but also the interplay between conflicting information sources, ultimately enhancing the reliability of multi-step question answering systems.
Search agents can fail dramatically when faced with unreliable evidence, revealing substantial performance disparities that traditional benchmarks overlook.
While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.