Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
17
By reallocating KV cache resources based on attention head specialization, SGD-KV cuts memory usage by 75% while handling contexts up to 1M tokens with state-of-the-art accuracy.