Search papers, labs, and topics across Lattice.
2
0
4
Scaling linear attention models with Sparse Delta Memory leads to significant improvements in long-context recall and reasoning without the computational burden of larger state sizes.
Simple prompting methods consistently outperform advanced dense supervision techniques, challenging the current assumptions in LLM training strategies.