Search papers, labs, and topics across Lattice.
1
0
2
7
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.