Search papers, labs, and topics across Lattice.
This paper introduces DRNOISE, a novel benchmark comprising 100 tasks designed to evaluate the performance of deep research agents in environments filled with misleading evidence. The study reveals that when agents encounter a plausible yet false document, their accuracy drops significantly by 66-88 percentage points, primarily due to a failure mode termed verification inertia, where agents retrieve accurate records but fail to reconcile them with conflicting claims. The findings underscore the necessity for deep research agents to not only retrieve information but also actively reconcile evidence to maintain sound evidential standards in open-web contexts.
A single misleading document can drastically reduce deep research agents' accuracy by up to 88%, exposing a critical vulnerability in their evidential reasoning capabilities.
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-looking false document is deliberately seeded into a searchable environment and offers a direct shortcut to a conflicting answer. We introduce DRNOISE, a 100-task benchmark for answer recovery under misleading evidence. Each task has a unique gold answer supported by two corroborating indirect record chains; the paired noisy condition adds one plausible document that states a conflicting answer directly. The benchmark spans ten families of evidence operations. Across agents with strong clean-task performance, this single intervention causes 66-88 percentage-point accuracy drops. Trace analyses identify verification inertia as the dominant failure mode: agents often retrieve truthful records but stop before completing and reconciling the evidence chain, instead deferring to the answer-like document. Generic verification prompts reduce but do not close this gap. The setting is especially relevant to open-web deployment, where plausible falsehoods arrive through ordinary-looking pages rather than explicit attacks. Reliable deep research therefore requires more than retrieval and citation; it requires active reconciliation of direct claims with record-level evidence.