Search papers, labs, and topics across Lattice.
4
0
5
3
Long-horizon post-training boosts software engineering performance by enhancing goal-directed execution behaviors, even when the training domain is unrelated.
Language models struggle to follow long, binding policies, with top configurations passing only 36.2% of trials in a rigorous benchmark simulating real-world tasks.
Despite advances in multimodal models, even the best can only answer 15% of realistic questions from professional PDFs, exposing significant gaps in current capabilities.
Today's best AI models are humbled by research-level math problems, scoring below 10% despite excelling at olympiad-style competitions.