Search papers, labs, and topics across Lattice.
Affiliation:, Baidu Inc.
4
0
5
17
Real-world complexities expose significant performance gaps in autonomous agents, revealing that even advanced LLMs struggle with task completion in dynamic environments.
Even the most advanced autonomous agents struggle with fundamental document manipulation tasks, revealing critical vulnerabilities in their operational capabilities.
Achieving new state-of-the-art scores in deep research benchmarks, DuMate-DeepResearch redefines the capabilities of multi-agent systems in tackling complex research tasks.
ToolMaze reveals that LLMs suffer a staggering 37% drop in recovery performance when faced with implicit semantic failures, highlighting a critical vulnerability in current models.