Search papers, labs, and topics across Lattice.
2
0
4
15
SeKV achieves a remarkable 5.9% performance boost in long-context LLM inference while slashing GPU memory usage by over half.
Reducing visual token usage by 46% while improving performance shows that CUAs can leverage more historical data effectively without overwhelming compute budgets.