Search papers, labs, and topics across Lattice.
2
0
3
3
TrustDABench reveals that even the best LLMs struggle with reliability and robustness in structured data analysis, achieving less than 25% accuracy in tracing evidence paths.
Span-level error localization can boost deep-research agent reliability by up to 30 percentage points, revealing critical insights into where agents go wrong.