Search papers, labs, and topics across Lattice.
Fudan University Shanghai Innovation Institute
1
0
4
Predicting the next token's KV entries can boost long-context LLM throughput by over 2.5 times without sacrificing latency or quality.