Search papers, labs, and topics across Lattice.
University of Science and Technology
1
0
3
Latent communication between LLMs was largely an illusion of interface shortcutting until Draft-KV, which routes genuine drafting KV caches to let a frozen 0.5B model leap from 37% to 78% MMLU-Redux accuracy by tapping an 8B model's latent thoughts.