Search papers, labs, and topics across Lattice.
Affiliation:
5
0
8
Even frontier LLMs collapse to sub-10% accuracy when facts in a long context must be dynamically updated or revoked, but training on executable state-machine simulations reliably repairs this state-tracking failure across diverse out-of-distribution tasks.
Explicit context compilation boosts LLM performance on in-context learning tasks, lifting accuracy from 15.4% to 21.4% on complex benchmarks.
Text world models can transform LLM-based agents from reactive responders into proactive planners, enhancing their performance in complex interactive tasks.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Force your VLMs to *show their work*: Saliency-R1 aligns model attention with human-annotated visual cues, boosting faithfulness and interpretability without extra compute.