Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Even frontier LLMs collapse to sub-10% accuracy when facts in a long context must be dynamically updated or revoked, but training on executable state-machine simulations reliably repairs this state-tracking failure across diverse out-of-distribution tasks.