Search papers, labs, and topics across Lattice.
4
0
5
5
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
LLMs exhibit a troubling tendency to prioritize code generation over genuine architectural understanding, revealing a critical gap in their evaluation metrics.
LLMs struggle to reliably apply skills, with performance varying significantly based on the agent harness used, challenging assumptions about their capability.
Ambiguity detection and effective clarification questioning are often neglected capabilities in LLM-powered search agents, leading to significant performance gaps.