Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Even frontier LLMs collapse to sub-10% accuracy when facts in a long context must be dynamically updated or revoked, but training on executable state-machine simulations reliably repairs this state-tracking failure across diverse out-of-distribution tasks.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.