Search papers, labs, and topics across Lattice.
Fudan University
2
0
4
Even the top-performing language models struggle with archive-grounded reasoning, achieving only 59.4% accuracy on a benchmark designed to test their agentic capabilities across diverse workplace documents.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.