Search papers, labs, and topics across Lattice.
2
0
3
2
ReCache achieves a staggering 92.43% reduction in KV-tensor memory while maintaining nearly identical performance in tool-augmented language models.
WIDE achieves a remarkable 55.1% performance boost at 50% sparsity, revolutionizing how LLMs can efficiently allocate computation at the token level.