Search papers, labs, and topics across Lattice.
Seoul National University
2
0
4
Recovering 87% of accuracy lost to KV eviction could redefine efficiency in reasoning tasks for large language models.
Diffusion language models can achieve up to 26x inference speedups with almost no accuracy loss, thanks to a clever entropy-based KV caching strategy that avoids costly full forward passes.