Search papers, labs, and topics across Lattice.
Seoul National University
1
0
2
Chaining low-rank KV corrections directly across sequential agent workflows eliminates redundant prefills, slashing peak GPU memory by 3.7x and cutting time-to-first-token in half without sacrificing model accuracy.