Search papers, labs, and topics across Lattice.
3
0
4
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
Even the most advanced LLMs struggle with consistent rubric verification, revealing substantial noise in scoring outputs across complex agentic scenarios.
LLMs struggle to simulate culturally nuanced emotional responses to bureaucratic processes, especially in Eastern cultures, suggesting current models lack the socio-cultural understanding needed for accurate policy simulation.