Search papers, labs, and topics across Lattice.
3
0
6
5
Current MLLMs struggle with active visual observation, scoring as low as 3.5% on tasks designed to test this critical cognitive function.
Despite impressive unit test pass rates, today's best LLMs rewrite code instead of precisely debugging it, achieving less than 45% edit precision even when explicitly instructed to minimize changes.
LLMs can reason through chains of thought 2.5x longer and solve more complex math problems by optimizing for the influence of each token on future reasoning steps.