Search papers, labs, and topics across Lattice.
3
0
4
3
Language models struggle to follow long, binding policies, with top configurations passing only 36.2% of trials in a rigorous benchmark simulating real-world tasks.
Despite advances in multimodal models, even the best can only answer 15% of realistic questions from professional PDFs, exposing significant gaps in current capabilities.
Today's best AI models are humbled by research-level math problems, scoring below 10% despite excelling at olympiad-style competitions.