Search papers, labs, and topics across Lattice.
Seoul National University
2
0
2
Multi-agent LoRA systems can slash time-to-first-token by 3.1脳 without sacrificing role specialization by decoupling shared base KV caches from precomputed low-rank adapter updates.
Chaining low-rank KV corrections directly across sequential agent workflows eliminates redundant prefills, slashing peak GPU memory by 3.7x and cutting time-to-first-token in half without sacrificing model accuracy.