Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
0
Complex scoring algorithms for KV cache compression are largely unnecessary: protecting the initial prompt and dropping reasoning tokens completely at random matches state-of-the-art accuracy with up to 43% higher serving throughput.