Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
PATTON, a PIM runtime that integrates production LLM serving engines with commodity PIM, achieves an average 1.95x speedup and 4.83x higher energy efficiency over evaluated baselines, requires no PIM processing-unit modifications, and maintains a KV cache hit rate comparable to the native GPU KV cache in vLLM.