Search papers, labs, and topics across Lattice.
Georgia Tech
2
0
3
Current scientific agents struggle to maintain a coherent narrative across evidence and calculations, with only 34.81% achieving strict accuracy in complex tasks.
Agents that ace long-context recall can still bomb when they need to use that memory to actually *do* something, revealing a critical flaw in how we currently evaluate memory in AI.