Search papers, labs, and topics across Lattice.
Peking University
2
0
4
Hypic slashes time-to-first-token by 2.45x and doubles throughput for hybrid-attention LLMs, all while preserving near-full accuracy.
Transforming the KV cache from a monolithic structure into a dynamic, head-aware system could revolutionize LLM serving efficiency and scalability.