Search papers, labs, and topics across Lattice.
3
0
5
0
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.
SwiftCache slashes latency by 69% and boosts context length nearly fourfold by enabling cross-model KV cache sharing on GPUs.
LLM scaling bottlenecks demand a shift towards cloud-native architectures and distributed systems, unlocking potential gains from serverless inference and quantum computing.