Search papers, labs, and topics across Lattice.
5
0
9
8
GPT-5.5 not only tops the leaderboard in policy evolution but also reveals critical insights into how agents can optimize performance through strategic feedback utilization.
Even the best LLMs struggle with Olympiad-level combinatorics, achieving only 65.4% on a benchmark designed to expose their reasoning limitations.
Current personal assistant agents struggle to anticipate and act on unstated user needs in long, complex workflows, revealing a critical gap between task completion and genuine proactivity.
LVLMs can maintain sharper visual focus during long-form generation by adding a lightweight, learnable memory module that bypasses attention dilution.
Reasoning SFT doesn't just memorize, it generalizes鈥攂ut only if you train it long enough, feed it good data, and use a capable model, and even then, reasoning gains come at the cost of safety.